Implicit Motion Handling for Video Camouflaged Object Detection

Xuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong, Mehrtash Harandi, Tom Drummond, Zongyuan Ge

Introduction

Video Camouflaged Object Detection (VCOD) is the task of discovering objects in a video that, appearance-wise, exhibit a great deal of similarity to the background scene. Despite enjoying wide applications (e.g., surveillance and security , autonomous driving , medical image segmentation , locust detection and robotics ), the problem of Camouflaged Object Detection (COD) is a daunting task as camouflaged objects are often indistinguishable to naked-eyes. This, in turn, has made VCOD a relatively under-explored problem in computer vision, as compared to several related problems such as video object detection (VOD) , video salient object detection (VSOD) , and video motion segmentation (VMS) .

In most computer vision tasks (e.g., instance segmentation , saliency detection ), it is assumed that objects have clear boundaries. This allows us to formulate the problem at the image level and even consider improvements if motion information is available. In contrast, object boundaries are ambiguous and indistinguishable when it comes to detecting camouflaged objects. This not only makes detection from images challenging, but also results in inaccurate estimation of optical flow and motion cues in videos .

The lack of clear boundaries means that the appearance of the camouflaged object resembles the background. This shows itself as two fundamental difficulties: 1) the object boundaries are often seamlessly blended into the background and is observable only when the object moves; 2) the object usually has repetitive textures similar to the environment; hence determining the movement of pixels across frames to estimate the motion (e.g., as done in optical flow) is erratic and erroneous. As the first difficulty, to successfully address VCOD, a neural network needs to effectively discover the nuances between the camouflaged object and the background with the help of motion information. Moreover, the motion information is inherently noisy and inaccurate according to the second difficulty, as shown in Figure 1. As such, employing VOD, VSOD, and VMS techniques may fail miserably if naively used or combined to address the VCOD problem.

In this work, we introduce SLT-Net, a new method to address VCOD that utilizes short-term dynamics and long-term temporal consistency to detect camouflaged objects in videos. Specifically, we employ a short-term dynamic module to implicitly capture the motion between consecutive frames. Rather than using optical flow to explicitly represent motions, we use a full-range correlation pyramid strategy to represent them implicitly. The primary motivation behind using a correlation pyramid is that even SOTA optical flow algorithms fail to estimate motions for camouflaged objects and their errors accumulate over the video’s duration. Also, it allows us to jointly optimize the motion estimation (implicitly) and the predictions with only the detection supervision. To provide a stable estimation, we further introduce a long-term refinement module to alleviate accumulated inaccuracies in the short-term dynamic module.

We realize the SLT-Net as a hybrid neural network with both transformer and CNN components. In particular, we use a transformer structure to encode features for constructing a correlation pyramid. Aside from its design flexibility, features extracted by the transformer contain global contextual information with long-range dependencies and less inductive bias , which we observe to be more distinguishable in estimating the motion.

While the correlation pyramid strategy can effectively capture motions for detecting camouflaged objects, it cannot scale gracefully to long video sequences due to its computational complexity. To solve this issue, we adopt a sequence-to-sequence model with a spatial-temporal transformer to refine the pair-wise prediction with long-term consistency across the videos as we empirically find it is more accurate than the standard ConvLSTM model .

Moreover, being a less-explored problem, large-scale datasets are not available to evaluate and benchmark VCOD systems. To promote new developments in this domain, we have curated a large-scale VCOD dataset based on the Moving Camouflaged Animals (MoCA) . The new dataset, or MoCA-Mask for short, contains 87 video sequences with 22,939 frames in total with pixel-wise ground truth masks. MoCA-Mask encapsulates a variety of challenges, such as complex backgrounds and tiny and well-camouflaged objects. We provide annotations, bounding boxes, and dense segmentation masks for every five frames for all the videos in the dataset. We also provide the first comprehensive benchmark for existing VCOD methods. In a nutshell, our contributions are as follows:

We propose a new VCOD framework that can effectively model short-term dynamics and long-term temporal consistency from videos, where the motion and the camouflage object segmentation can be jointly optimized through a single optimization target.

We collect the first large-scale VCOD dataset, the MoCA-Mask dataset, to promote developments in VCOD as well as a comprehensive VCOD benchmark to facilitate research in VCOD.

We set a new state-of-the-art on the VCOD task, outperforming a previous SOTA method by 9.88%9.88\%.

Related Work

COD. Without any prior, even humans can easily miss camouflaged objects. However, once informed that a camouflaged object exists in an image, we can carefully scan the entire image to identify it. Inspired by this fact, ANet incorporated classification stream as the awareness of camouflaged objects and segmentation stream. Sharing a similar idea, SINet and PFNet addressed the problem by first positioning coarse camouflaged objects and then refining it by segmentation. SINet-v2 extended this idea by incorporating the reverse guidance before learning complementary regions. MGL incorporated edge details into the segmentation stream via two graph-based modules. By modeling the conspicuousness of camouflaged objects against backgrounds, Lv et al. introduced two new tasks, namely camouflaged object ranking and camouflaged object localization, along with relabeled NC4K dataset.

VSOD. To detect salient objects in videos, DLVS introduced fully convolutional networks for pixel-wise saliency prediction. DSR3 exploited an end-to-end 3D neural network to produce video sequences, which incorporates 3D CNN modules combined with recurrent refinement units to predict saliency maps. To better learn temporal information over frames, following works considered SpatioTemporal CRF , pyramid dilated convLSTM in the design of their networks. FGRN , RCRNet adopted extra flow-guided networks to improve temporal coherence. Later, SSAV specifically focused on the saliency shift phenomenon and established a comprehensive benchmark for VSOD. FSNet leveraged the mutual constraints of appearance and motion cues, demonstrating superior performances to many existing methods.

VMS. The task of VMS focuses on discovering moving objects in videos. Traditional methods usually address this problem by extracting motion boundaries in the flow field and then refining the initial estimate with appearance features , or combining motion and appearance cues by a fusion architecture . Another line of work explicitly leverages optical flow as the input to train a CNN-based network and generate pixel-level motion labels based on supervised learning or in an unsupervised manner .

VCOD. Different from VMS, visual cues of camouflage objects are considered less effective than motion cues. Prior works mainly relied on homography or optical flows to detect motion patterns. Bidau et al.proposed to segment moving objects from the environment by approximating different motion models computed from dense optical flow . In particular, in authors proposed a two-step segmentation algorithm, which first compensated for the camera rotation and then segmented the angle of the optical flow into objects and the background. Although each motion model is updated with optical flow orientations over time, the initial motion is heuristic. In , authors used a network to segment the angle field rather than raw optical flow. proposed a video registration and motion segmentation framework, along with a larger camouflaged dataset (MoCA) labeled by bounding boxes for every five frames. The explicit alignment method by optical flow builds spatial correspondence between neighboring frames. However, the optical flow estimation may not be accurate enough to support effective alignment, particularly in dynamic scenes with fast object motions.

Proposed Framework

To train the SLT-Net, we adopt a two-stage strategy. We first train the short-term detection module using pixel-wise annotations only. Once the model converges, we attach the long-term refinement module to the SLT-Net and train the whole model while fixing the short-term detection module.

2 Short-term Architecture

We illustrate our short-term architecture in Figure 3. It takes two consecutive frames as input from a video and predicts a binary mask of the reference frame. Our model consists of three main modules: (1) Transformer Encoder for feature extraction; (2) Short-term Correlation Pyramid for capturing short-term dynamics; and (3) CNN Decoder to predict the short-term segmentation. Below we describe the details of each module.

1. Transformer Encoder. We adopt a Siamese structure with the pyramid vision transformer (PVT) to extract features from two consecutive frames. The encoder consists of four stages that generate feature maps at four different scales. All stages share a similar structure, including a patch embedding layer and transformer blocks. The sizes of the features at each stage are Ci×H/2i+1×W/2i+1C_{i}\times H/2^{i+1}\times W/2^{i+1}, i∈{1,2,3,4}i\in\{1,2,3,4\}, where the H,W,CH,W,C represent the height, the width and the channels. We set C=32C=32 in our experiments. Following , we adapt three texture enhanced modules (TEM) for the features from the last three stages. To attain more discriminative feature representations, each TEM includes four parallel residual branches.

2. Short-term Correlation Pyramid. Prior works (e.g., ) explicitly incorporate motion by taking optical flow from consecutive frames as the inputs into a deep network. However, the inaccurate optical flow may result in error accumulation at subsequent predictions. If we would like to optimize the optical flow module with the segmentation module jointly, the ground truth of optical flow is required. To solve this issue, inspired by , we propose a correlation pyramid to capture motion information implicitly. As shown in Figure 3, the CNN decoder directly takes the correlation pyramid as its only input. It means the network can only estimate correct segmentation with correct motion estimation. Also, since the features used to form the correlation pyramid will be updated with the segmentation ground truth, we can use the segmentation ground truth to optimize motion estimations and detection results jointly.

with cc being the index along the channel dimension of frame features. With all neighboring features are paired up with correlations, we can find correspondences at a global scale. To reduce the computational complexity, we downsample the adjacent frame by max-pooling over features while keeping the resolution of the reference frame. This design helps the model to learn multi-scale displacement while maintaining high-resolution image details.

Next, we normalize the feature correlation volume C(It, It+1)xyuv\mathbf{C}(\mathbf{I}_{t},\,\mathbf{I}_{t+1})_{xyuv} along the last two dimensions uvuv over their sum, as they represent the correspondence between the reference and downsampled neighboring feature frame in all the spatial position. The normalized correlation volume is computed as follows:

Figure 4 only shows a correlation on one scale. To make the network learn more detailed information, we construct a correlation pyramid {Ci},i∈{2,3,4}\{\mathbf{C}^{i}\},i\in\{2,3,4\} by incorporating the extracted multi-scale features from the transformer encoder (See details in supplementary materials (Supp) ).

Learning Strategy. We train the short-term training stage by minimizing the loss below:

The weighted cross-entropy loss Lcew\mathcal{L}^{w}_{ce} increases the weights of hard pixels to emphasize their importance. The weighted intersection-over-union loss Liouw\mathcal{L}^{w}_{iou} pays more attention to hard pixels rather than assigning all pixels with equal weights. Readers could refer to prior work to find more details regarding the definitions of these two loss functions.

3 Long-term Consistency Architecture

There are two kinds of seq-to-seq modeling architecture: one uses convLSTM to model the temporal information, and the other uses a transformer-based seq-to-seq modeling network. We implement both architectures and compare their results in Section 4.4. We empirically find that using the transformer structure can lead to better results, so we select it as our seq-to-seq modeling network to enforce the long-term consistency.

We show the details of the seq-to-seq modeling network on the right side of Figure 5. For each target pixel, to reduce the complexity for building a dense spatial-temporal affinity matrix, we select a fixed number of relevance measuring blocks to construct the affinity matrix within a constrained neighborhood of it. We apply the hybrid loss during the training:

where Le\mathcal{L}_{e} is the Enhanced-alignment loss, the hybrid loss can guide the network to learn pixel-, object- and image-level features.

Experiments

This section performs a thorough evaluation of our proposed framework on the CAD dataset and our proposed MoCA-Mask dataset. We also provide a comprehensive VCOD benchmark to facilitate the research of VCOD.

COD10K. We pre-train all still image-based methods as well as the encoder of the video-based methods on COD10K . It is currently the largest COD dataset which consists of 5,066 camouflaged images (3,040 for training, 2,026 for testing), and is divided into five super-classes and 69 sub-classes. This dataset also provides high-quality annotation, reaching the level of matting.

CAD. Camouflaged Animal Dataset (CAD) is a small set of camouflaged animals, first introduced by . It includes nine short video sequences in total that were extracted from YouTube videos and accompanying hand-labeled ground-truth masks on every 5th5^{th} frame. We also provide pseudo GT masks by a bidirectional consistency check strategy to enable future studies on this dataset.

MoCA-Mask. The original Moving Camouflaged Animals (MoCA) Dataset includes 37K frames from 141 YouTube Video sequences with resolution and sampling rate of 720×1280720\times 1280 and 24fps in the majority of cases. The dataset covers 67 types of animals moving in natural scenes, but some are not camouflaged animals. Also, the ground truth of the original dataset is bounding boxes rather than dense segmentation masks, which makes it hard to evaluate the VCOD segmentation performance. To this end, we reorganize the dataset as MoCA-Mask and build a comprehensive benchmark with more comprehensive evaluation criteria. The modifications could be found in Supp.

2 Benchmarks

Metrics. We adopt the following evaluation metrics to measure the pixel-wise masks: (1) MAE (MM), which assesses the pixel-level accuracy between prediction and labeled masks. (2) Enhanced-alignment measure (EϕE_{\phi}) , which simultaneously evaluates the pixel-level matching and image-level statistics. This metric is naturally suited for assessing the overall and localized accuracy of the camouflaged object detection results. Note that we report mean EϕE_{\phi} in the experiments. (3) S-measure (SαS_{\alpha}) , which evaluates region-aware and object-aware structural similarity. (4) Weighted F-measure FβwF_{\beta}^{w} can provide more reliable evaluation results than the traditional FβF_{\beta}. (5) mean Dice, which measures the similarity between two sets of data. (6) meanIoU, which measures the overlap between two masks.

Baseline. We select nine cutting-edge baselines, including I. six image based methods i.e., EGNet , BASNet , CPD , PraNet , SINet , SINet-v2 , and II. three video based methods, i.e., PNS-Net , RCRNet , and MotionGroup . Please refer to the Supp for the implementation details.

Settings. We compare our method primarily with the top-performing single image and video baselines. As network architectures, input resolution, modality, pre-processing, and post-processing are all different, we try our best to conduct the comparison as fairly as possible. For single image baselines, we adopt the same data pre-processing as for all the compared methods. Specifically, the input images are resized to 352×352352\times 352, after random flip, random rotation, and color enhance augmentation. In the training phase, we apply random pepper noise on the GT images. As EGNet requires extra edge/boundary information for training, we adopt the same pre-processing techniques in their paper to obtain the edge maps. This extra information could also be found in our reorganized version of the MoCA-Mask dataset.

Most of the video approaches, e.g., PNS-Net , RCRNet , employ a multistage training pipeline. The model is pre-trained using still image datasets and then equipped with temporal modules to process video datasets. We follow this training strategy and pre-train all methods on the COD10K training set, except MotionGroup which does not have a static model. Also, per our practical experience, loading pre-trained weights on the COD10K dataset could further improve the model performance on MoCA-Mask. Compared with the COD10K image dataset, the video dataset MoCA-Mask is more challenging due to the camera motions, blurring images, small ratio of animals, and their tiny body structures, such as slim torso/limbs. In some video sequences, the animals make up a tiny proportion of the entire frame, which makes them extremely hard to be identified (see, for example, ibex in Figure 6). Based upon the considerations above, we provide the results based on the following setting: (a) Training the models on COD10K; (b) Fine-tuning the models on MoCA-Mask, with pre-trained weights on COD10K; (c) Evaluate the models on the whole CAD, the test set of MoCA-Mask.

3 Results

Performance on MoCA-Mask. In Table 1, our approach outperforms all the studied methods by a significant margin, notably by 9.88%9.88\% on SαS_{\alpha} over the best one in this evaluation, RCRNet , and 92.97%92.97\% on FβwF_{\beta}^{w} metric over SINet . We also provide the qualitative comparisons of our method and other baselines in Figure 6. Our model can accurately locate and segment camouflaged objects in many challenging situations, such as objects with the tinny torso or complex appearance textures, blur, or abrupt motions. We provide more details, i.e., per-sequence quantitative and qualitative results in the Supp, to illustrate the consistent success over the consecutive frames.

Performance on CAD. In Table 2, we assess different approaches by studying their cross-dataset generalization on the CAD dataset. Again, the proposed network obtains the best performance in terms of all six evaluation metrics, further demonstrating its robustness. As shown in Figure 7, our model achieves sharper boundaries with more fine-grained visual details. This benefits from constructing pixel-level correlation pairs in the feature space.

4 Ablation Studies

We perform ablation studies on the MoCA-Mask dataset. In particular, we look into functionality analysis for our short-term and long-term modules, the choice of sequence-to-sequence model, and our pseudo masks.

Short-term and Long-term Modules. We evaluate the effectiveness of our short-term and long-term modules in two aspects. We first perform an ablation study on the short-term and the long-term modules on the MoCA-Mask dataset and show the results in Table 3. By adding the short-term module, our performance is improved by 2.16% on SαS_{\alpha}, 6.06% on FβwF_{\beta}^{w}, 2.41% on EϕE_{\phi}, 16.00% on MM, 4.53% on mDic, and 4.84% on mIoU. By adding the long-term module, we further improve our performance by 2% on FβwF_{\beta}^{w}, while a slight drop 0.91% on SαS_{\alpha}.

We then swap the encoder of a SOTA VSOD method RCRNet with our transformer based encoder to compare the effectiveness of the temporal information handling strategies between ours and the RCRNet in Table 4. In terms of its spatiotemporal coherence model, it shows both positive and negative gains on the evaluated metric, i.e., 1.51% on SαS_{\alpha}, -0.97% on FβwF_{\beta}^{w}, -0.16% on EϕE_{\phi}, 6.98% on MM.

Transformer v.s. ConvLSTM. We evaluate two different approaches for constructing long-term architecture, namely transformer based model, and ConvLSTM based model. For the latter ConvLSTM network variant, we adopt a sequence model proposed in but modify the original VGG-style network for the CNN encoder and decoder with our transformer-style backbone network. From the Table 5, we can observe that the transformer variant is more accurate than the ConvLSTM model in all four metrics, with a much smaller number of parameters.

Pseudo Masks. As shown in Table 1, although the generated pseudo labels contain some noises, they can improve the performance of video approaches as they can leverage temporal information to suppress the label noises. For still image baselines, almost all of them are seriously effected by the label noises, leading to worse performance than the one without pseudo labels. It also proves that the motion estimation error can not be overlooked in the VCOD problem and we should jointly optimize it with the segmentation error for a better performance.

Trained from scratch on MoCA-Mask. For the sake of completeness, we provide the accuracy of our network with/without pre-trained weights in Table 6. It shows that the gap between the train-from-scratch and the pre-trained model is minor, i.e., only a slight drop 0.15% on SαS_{\alpha}.

Generalization. Our model can be applied to the more general video object detection problem, such as video instance segmentation. Except a detailed comparison with MG in Table 1, we compare with on DAVIS16 (Table 7) and demonstrate superiority of our method.

Conclusion

We presented a new SLT-Net framework for learning to segment camouflaged objects in a video. Specifically, we proposed a short-term module to implicitly capture motions between consecutive frames which allows us to learn motion estimation and segmentation in a single optimization target. We also proposed a long-term module with a sequence-to-sequence transformer to enforce temporal consistency in video sequence. To promote the development of this field, we rebuild a new dataset called MoCA-Mask with 87 high-quality video sequences, including 22,939 frames in total. It is the largest-scale pixel-level annotated dataset that allows object-level benchmark in video camouflaged object detection (VCOD). Compared with existing state-of-the-art baselines, our proposed network achieves fascinating results on two VCOD benchmarks.

Broader Impact. Camouflaged object can be used to detect and protect rare animal species, prevent wildlife trafficking, medical applications (e.g., detecting polyp or lung infection) and search-and-rescue work to name a few. Please note that our MoCA-Mask dataset does not contain any military or sensitive scenes. Aside from its important use-cases as mentioned above, our paper takes a solid step into understanding video contents when motion information is noisy.

References

Supplementary Material

Semi-supervised Training Procedure

As the annotations are provided in the form of dense segmentation masks for every five frames, we adopt a bi-directional consistency check strategy to generate pseudo masks for unlabelled frames. Given five consecutive frames {It,It+1,It+2,It+3,It+4}\{\mathbf{I}_{t},\mathbf{I}_{t+1},\mathbf{I}_{t+2},\mathbf{I}_{t+3},\mathbf{I}_{t+4}\} and labelled ground-truth gtt\mathbf{gt}_{t}, we first estimate forward and backward optical flow fields between frame It\mathbf{I}_{t} and It+n,n∈\mathbf{I}_{t+n},n\in. Then we can produce the warped ground-truth gt^t+n\hat{\mathbf{gt}}_{t+n} with the inverse warping from ground-truth gtt\mathbf{gt}_{t}.

1.Flow Estimation. We take the ground-truth mask of the reference frame It\mathbf{I}_{t} as an example, to generate pseudo ground-truth of its immediate following frame It+1\mathbf{I}_{t+1}. The optical flow estimation moduleIn practice, we make use of RAFT to obtain the optical flow. O\mathcal{O} takes It\mathbf{I}_{t} and It+1\mathbf{I}_{t+1} and predicts the optical flow field:

where ut,t+1x\mathbf{u}_{t,t+1}^{x} and ut,t+1y\mathbf{u}_{t,t+1}^{y} denote the x, yx,\,y components of the estimated flow field, respectively. The flow field maps each pixel (x, y)(x,\,y) in It+1{\mathbf{I}}_{t+1} to its corresponding coordinates (x′, y′)=(x+ut,t+1x(x), y+ut,t+1y(y))(x^{\prime},\,y^{\prime})=(x+\mathbf{u}_{t,t+1}^{x}(x),\,y+\mathbf{u}_{t,t+1}^{y}(y)) in It{\mathbf{I}}_{t}.

2.Forward/Backward Pseudo Labeling. Given the forward optical flow sequences (flowt,flowt+n),n∈1,2,3,4(\mathbf{flow}_{t},\mathbf{flow}_{t+n}),n\in{1,2,3,4}, we can obtain the aligned neighboring frame gt^t+n\hat{\mathbf{gt}}_{t+n} by a warping interpolation on gtt\mathbf{gt}_{t} using the mapped coordinates. After repeating the explicit alignment step for the preceding frame, we acquire the sequence of warped input frames {gtt,gt^t+1,gt^t+2,gt^t+3,gt^t+4}\{\mathbf{gt}_{t},\hat{\mathbf{gt}}_{t+1},\hat{\mathbf{gt}}_{t+2},\hat{\mathbf{gt}}_{t+3},\hat{\mathbf{gt}}_{t+4}\}. The backward pseudo ground-truth sequences are obtained by performing warping ground-truth masks with backward optical flows in the reverse order.

3.Bidirectional Consistency Check. To identify valid masks, we adopt forward-backward consistency check to eliminate inconsistent regions. Under the forward-backward consistency assumption , traversing flow vector forward and then backward should arrive at the same position. We mark pixels as invalid whenever this constraint is violated. As shown in Figure 9, the invalid regions emphasized by the orange boxes are marked as background.

Training Details

We implement both long-term and short-term architecture in PyTorch. The input images are resized to 352×352352\times 352. We train the short-term architecture with a batch size of 8 on an NVIDIA V100 GPU and use Adam optimizer with initial learning rate of 1e-4, decreasing every 50k iterations. For the long-term optimization, our model takes 10 frames as the input at one time with the frame sampling rate 1. For our pseudo ground-truth generation, we exploit RAFT as the optical flow estimation module and pre-trained weights on Sintel dataset .

Data Curation

Remove Invalid Scenes. We first select and exclude scenarios in that animals are obvious and easy to identify from the background at our first glance. After cleaning the dataset, our new subset includes 87 video sequences, 22,939 frames in total.

Segmentation Masks. For annotations, we further provide accurate human-labeled segmentation masks for every five frames. Thus our GT consists of two formats, that is 4,691 bounding box annotations as well as 4,691 pixel-level masks.

Pseudo Masks. We use a bidirectional optical flow-based strategy to generate the pseudo GT masks, refer to the SM. Note that these pseudo masks still contain motion estimation errors, requiring algorithms to have the capability to handle noise labels when using them.

Dataset Split. The whole dataset is split into 71 sequences, 19,313 frames for training, and 16 sequences, 3,626 frames selected for testing. The summary of each sub-sequence distribution could be found in Fig. 11.