TokenCut: Segmenting Objects in Images and Videos with Self-supervised Transformer and Normalized Cut
Yangtao Wang, Xi Shen, Yuan Yuan, Yuming Du, Maomao Li, Shell Xu Hu, James L Crowley, Dominique Vaufreydaz
Introduction
Detecting and segmenting salient objects in an image or video are fundamental problems in computer vision with applications in real-world vision systems for robotics, autonomous driving, traffic monitoring, manufacturing, and embodied artificial intelligence . However, current approaches rely on supervised learning requiring large data sets of high-quality, annotated training data . The high cost of this approach becomes even more apparent when using transfer learning to adapt a pre-trained object detector to a new application domain. Researchers have attempted to overcome this barrier using active learning , semi-supervised learning , and weakly-supervised learning with limited results. In this paper, we report on results of an effort to use features provided by transformers trained with self-supervised learning, obviating the need for expensive annotated training data.
Vision transformers trained with self-supervised learning , such as DINO and MAE have been shown to outperform supervised training on downstream tasks. In particular, the attention maps associated with patches typically contain meaningful semantic information (Fig. 1(a)). For example, experiments with DINO indicate that the attention maps of the class token highlight salient object regions. However, such attention maps are noisy and cannot be directly used to detect or segment objects.
The authors of LOST have shown that the learned features from DINO can be used to build a graph and segment objects using the inverse degrees of nodes. Specifically, LOST employs a heuristic seed expansion strategy to accommodate noise and detect a single bounding box for a foreground object. We have investigated whether such learned features can be used with a graph-based approach to detect and segment salient objects in images and videos (Fig. 1(b)), formulating the segmentation problem using the classic normalised cut algorithm (Ncut) .
In this paper we describe TokenCut, a unified graph-based approach for image and video segmentation using features provided by self-supervised learning. The processing pipeline for this approach, illustrated in Fig. 2, is composed of three steps: 1) graph construction, 2) graph cut, 3) edge refinement. In the graph construction step, the algorithm uses image patches as nodes and uses features provided by self-supervised learning to describe the similarity between pairs of nodes. For images, edges are labeled with a score for the similarity of patches based on learned features for RGB appearance. For videos, edge labels combine similarities of learned features for RGB appearance and optical flow.
To cut the graph, we rely on the classic normalized cut (Ncut) algorithm to group self-similar regions and delimit the salient objects. We solve the graph-cut problem using spectral clustering with generalized eigen-decomposition. The second smallest eigenvector provides a cutting solution indicating the likelihood that a token belongs to a foreground object, which allows us to design a simple post-processing step to obtain a foreground mask. We also show that standard algorithmns for edge-aware refinement, such as Conditional Random Field (CRF) and Bilateral Solver (BS) can be used to refine the masks for detailed object boundary detection. This approach can be considered as a run-time adaptation method, because the model can be used to process an image or video without the need to retrain the model.
Despite its simplicity, TokenCut significantly improves unsupervised saliency detection in images. Specifically, it achieves 77.7%, 62.8%. 61.9% mIoU on the ECSSD , DUTS and DUT-OMRON respectively, and outperforms the previous state-of-the-art by a margin of 4.4%, 5.6% and 5.2%. For unsupervised video segmentation, TokenCut achieves competitive results on DAVIS , FBMS , SegTV2 . Additionally, TokenCut also obtains important improvement on unsupervised object discovery. For example, TokenCut outperforms DSS , which is a concurrent work, by a margin of 6.1%, 5.7%, and 2.6% respectively on the VOC07 , VOC12 , COCO20K .
In summary, the main contributions of this paper are as follows:
We describe TokenCut , a simple and unified approach to segment objects in images and videos that does not require human annotations for training. The implementation is available at https://www.m-psi.fr/Papers/TokenCut2022/. An online demo is accessible at https://huggingface.co/spaces/yangtaowang/TokenCut (last access May 2023).
We show that TokenCut significantly outperforms previous state-of-the-art methods unsupervised saliency detection and unsupervised object discovery on images. As a training-free method, TokenCut achieves competitive performance on unsupervised video segmentation compared to the state-of-the-art methods.
We provide a detailed analysis on the TokenCut to validate the design of the proposed approach.
Related Work
ViT has shown that the transformer architecture can be effective for computer vision tasks using supervised learning. Recently, many variants of ViT have been proposed to learn image encoders in a self-supervised manner. MoCo-V3 demonstrates that using contrastive learning on ViT can achieve strong results. DINO shows that transformers can be trained with self-distillation loss and shows that the features learn by ViT contain explicit information useful for image semantic segmentation. Inspired by BERT, several approaches learn by missing token replacement, masking some tokens from the input input and learning to recover the missing tokens in the output.
Given a group of images, unsupervised object discovery seeks to discover and delimit similar objects that appear in multiple images. Early research formulated the problem using an hypothesis about the frequency of object occurrences. Other researchers formulated object detection as an optimization problem over bounding box proposals or as a ranking problem .
Recently, LOST significantly improved the state-of-the-art for unsupervised object discovery. LOST extracts features using a self-supervised transformer based on DINO and designs a heuristic seed expansion strategy to obtain a single object region. A concurrent work DSS designs a weighted graph over patches using self-supervised transformer features as well as a KNN based image matting algorithm using a color affinity matrix. The eigen-decomposition of the affinity matrix is computed to obtain a coarse object mask. Following LOST , DSS also assumes that the foreground object occupies a smaller region than the background. While DSS has a similar eigen-attention map as TokenCut, DSS is less able to detect large objects.
As with LOST, TokenCut also uses features obtained with self-supervised learning. However, rather than relying on the attention map of some specific nodes, TokenCut forms a fully connected graph of image tokens, with edges labeled with a similarity score between tokens based on transformer features. The classical Ncut algorithm is then used to detect and segment image objects.
Unsupervised saliency detection seeks to segment a salient object within an image. In this paper, we show that incorporating a simple post-processing step into TokenCut for unsupervised object discovery can provide a strong baseline method for unsupervised saliency detection. Earlier works on Unsupervised saliency detection use techniques such as color contrast , background priors , or super-pixels . More recently, unsupervised deep models have been used for saliency detection using noisy pseudo-labels generated from different handcrafted saliency methods. shows that unsupervised GANs can differentiate between foreground and background pixels and generate high-quality saliency masks. SelfMask use a spectral clustering method with self-supervised features to group pixels into a set of candidate clusters. SelfMask is trained by selecting the salient masks from the set of spectral clusters as pseudo-masks for supervision using cluster voting scheme.
Given an unlabeled video, unsupervised video segmentation aims to generate pixel-level masks for the object of interest in the video. Prior works segment objects by selecting super-pixels , learning flattened 3D object representations , constructing an adversarial network to mask a region such that the model can predict the optical flow of the masked region , or reconstructing the optical flow in a self-supervised manner , etc. DyStaB first partitions the motion field by minimizing the temporal consistent mutual information and then uses the segments to learn the object detector, in which the models are jointly trained with a bootstrapping strategy. The deformable sprites method (DeSprite) trains a video auto-encoder to segment the object of interest by decomposing the video into layers of persistent motion groups. In contrast to these methods , our proposed method does not require prior training on videos. Compared with methods that do not train on videos, our method achieves superior performance.
Approach: TokenCut
In this section, we present TokenCut, a unified algorithm that can be used to segment salient objects in an image or moving objects in a video. Our approach, illustrated in Fig. 2, is based on a graph where the nodes are visual patches from either an image or a sequence of frames, and the edges are similarities between the features of the nodes based on the features provided by a visual transformer trained with self-supervised learning.
This section is organised as follows: we first briefly review vision transformers and the Normalized Cut algorithm in Section 3.1.1 and Section 3.1.2. We then describe the TokenCut algorithm for object detection and segmentation in images and videos in Section 3.2.
The Vision Transformer has been proposed in . The key idea is to process an image with transformer architectures using non-overlapping patches as tokens. For an image with size , a vision transformer takes non-overlapping image patches as inputs, resulting in patches. Each patch is used as a token, described by a vector of numerical features that provide an embedding. An extra learnable token, denoted as a class token , is used to represent the aggregated information of the entire set of patches. A positional encoding is added to token and the set of patch tokens, and the resulting vector is fed to a standard Vision Transformer with self-attention and layer normalization .
The Vision Transformer is composed of several stacked layers of encoders, each with feed-forward networks and multiple attention heads for self-attention, paralleled with skip connections. For the TokenCut algorithm, we use the Vision Transformer, trained with self-supervised learning. We extract latent features from the final layer as the input features for TokenCut.
1.2 Normalized Cut (Ncut)
Given a graph = (, ), where and are sets of nodes and edges respectively. is the similarity matrix with as the edge between the i-node and the j-th node . Ncut is proposed to partition the graph into two disjoint sets and . Different to standard graph cut, Ncut criterion considers both the total dissimilarity between and as well as the total similarity within and . Precisely, we seek to minimize the Ncut energy :
where measures the degree of similarity between two sets. , and is the total connection from nodes in to all the nodes in the graph.
As shown by , the equivalent form of optimization problem in Eqn 1 can be expressed as:
with the condition of , where satisfies , where is a diagonal matrix with on its diagonal.
Taking , Eqn 2 can be rewritten as:
Indicating in , the formulation in Eqn 3 is equivalent to the Rayleigh quotient , which is equivalent to solve , where is the Laplacian matrix and known to be positive semidefinite . Therefore is an eigenvector associated to the smallest eigenvalue . According to the Rayleigh quotient , the second smallest eigenvector is perpendicular to the smallest one () and can be used to minimize the energy in Eqn 3,
Taking ,
Thus, the second smallest eigenvector of the generalized eigensystem provides a solution to the Ncut problem.
2 The TokenCut Algorithm
The TokenCut algorithm consists of three steps: (a) Graph Construction, (b) Graph Cut, (c) Edge Refinement. An overview of the algorithm is shown in Fig. 2.
As described in Section 3.1.2, TokenCut operates on a fully connected undirected graph = (, ), where represents the feature vectors of the node . Each patch is linked to other patches by labeled edges, . Edge labels represent a similarity score .
where is a hyper-parameter and is the cosine similarity between features. is a small value to assure a fully connected graph. Note that the spatial location information has been implicitly included in the features, which is achieved by positional encoding in the transformer.
As with images, videos are presented as a fully connected graph where the nodes are visual patches and the edges are labeled with the similarity between patches. However, for videos, similarity includes a score based on both RGB appearance and a RGB representation of optical flow computed between consecutive frames . The algorithm extracts a sequence of feature vectors using a vision transformer as described in Section 3.1.1. Let and denote the feature of i-th image patch and flow patch respectively. Edges are labeled with the average over the similarities between image feature and flow features, expressed as:
Image feature provide segmentation using appearance while flow features focus on segmentation with motion. We provide a full analysis on the definition of edges in Section 4.5.
2.2 Graph Cut
The Ncut algorithm is used to partition the fully connected graph. Ncut computes the second smallest eigenvector of the generalized eigensystem, as described in Section 3.1.2 to highlight salient objects. We refer to this eigenvector as a measure for “eigen-attention”, and provide visualizations of the attention map provided by this vector in Section 4. TokenCut uses eigen-attention to bi-partition the graph, determines which partition belongs to the foreground and then determines the nodes that belong to the each object region.
To partition the nodes into two disjoint sets, TokenCut uses the average value of the second smallest eigenvector to cut the graph . Formally, and . Note that, we also explored the use of classical clustering algorithms, such as K-means and EM, to cluster the second smallest eigenvector into 2 partitions. The comparison is available in Section 4.5. Our experimens show that the average value generally provides better results.
Given the two disjoint sets of nodes, TokenCut selects the partition with the maximum absolute value as the foreground. Intuitively, the foreground object should be salient and thus less connected to the entire graph. In other words, if belongs to the foreground while is the background token. Therefore, the eigenvector of the foreground object should have a larger absolute value than the background region.
In images, we are interested in segmenting a single object. However, the foreground can contain more than one salient object region. TokenCut selects the connected component in the foreground containing the maximum absolute value as the detected object. In videos, as the goal is to segment objects based on both motion and appearance, TokenCut takes the entire foreground region as the final output.
2.3 Edge Refinement
The graph cut algorithm provides coarse masks of object regions due to the large size of transformer patches. The boundaries of such masks can be easily refined using standard edge refinement technique. We have experimented with off-the-shelf edge-aware post-processing techniques such as Bilateral Solver (BS), Conditional Random Field (CRF) on top of the obtained coarse mask to generate more precise boundaries for the mask. We have found that CRF usually provides the best results.
Experiments
We evaluated the suitability of TokenCut for three tasks: unsupervised single object discovery, unsupervised saliency detection and unsupervised video segmentation. We present implementation details in Section 4.1. The results of unsupervised single object discovery are shown in Section 4.2. The results for unsupervised saliency detection are presented in Section 4.3, and results for unsupervised video segmentation in Section 4.4. We provide ablation studies in Section 4.5.
For our experiments, we use the ViT-S/16 model trained with self-distillation loss (DINO) to extract features of patches. Following , we employ the key features of the last layer as the input features . Ablations on different features and ViT backbones are provided in Tab. 5. We set for all image datasets and for video datasets. The selection of is discussed in Section 4.5.
In terms of running time, our implementation takes approximately 0.32 seconds to detect a bounding box for a salient object region in a single image with resolution 480 480 using a single GPU QUADRO RTX 8000. Obtainaing a coarse mask from 20 frames of video with 320 x 576 resolution, requires an average of 30 seconds with standard deviation of around 4.5 seconds. Edge refinement the takes an additional 16.4 seconds on average with a standard deviation 1.4 seconds. As with single-frame graphs, the same video takes 0.93 seconds in average to obtain the coarse mask with standard deviation of 0.17 for all frames. The post processing step cost 16.1 seconds with standard deviation of 1.4. For n tokens, the algorithmic complexity for building such a graph is . Thus the average processing time grows with the square of the number of frames in the video.
To generate optical flow, we use two different approaches: RAFT and ARFlow . The first one is supervised and the second one is self-supervised. We extract the optical flow at the original resolution of the image pairs, with the frame gaps n = 1 for DAVIS and SegTV2 dataset. For FBMS we use n = 3 to compensate for the much slower rate of motion. This improves the optical flow quality as small pixel-level motions are hard to detect using off-the-shelf methods. Optical flow features are encoded as RGB values, using standard techniques for visualization of optical flow . This allows us to directly use the pre-trained self-supervised transformers with optical flow encoded as RGB. Because of limits on available computational resources, we construct the video graph with a maximum of 90 frames on the DAVIS dataset. For videos longer than 90 frames, it is possible to aggregate results using non-overlapping subgraphs with maximum video frames of 90.
2 Unsupervised Single Object Discovery
TokenCut has been evaluated on three commonly used benchmarks for unsupervised single object discovery: VOC07 , VOC12 and COCO20K . VOC07 and VOC12 contain 5011 and 11540 images respectively which belong to 20 categories. COCO20K consists of 19817 randomly chosen images from the COCO2014 dataset . VOC07 and VOC12 are commonly used to evaluate unsupervised object discovery . COCO20K is a popular benchmark for a large scale evaluation .
In line with previous research , we report performance using the CorLoc metric for precise localization. We use a single predicted bounding box for each image. For target images, CorLoc is 1.0 if the intersection over union (IoU) score between the predicted bounding box and the ground truth bounding boxes is superior to 0.5.
We evaluate the CorLoc scores in comparison with previous state-of-the-art single object discovery methods on VOC07, VOC12, and COCO20K datasets. These methods can be roughly divided into two groups according to whether the model uses information from the entire dataset or explores inter-image similarities. Because of the quadratic complexity of region comparison among images, models with inter-image similarities are generally difficult to scale to larger datasets. The selective search , edge boxes , LOST and TokenCut do not require inter-image similarities and are thus much more efficient. As shown in the Tab. 1, TokenCut consistently outperforms all the previous methods on all the datasets by a large margin. Particularly, TokenCut ouperforms DSS by 6.1%, 5.7% and 2.6% for VOC07, VOC12 and COCO20K respectively using the same ViT-S/16 features.
We also list a set of results that includes using a second stage unsupervised training strategy to boost the performance. This is referred to as Class-Agnostic Detection (CAD) and proposed in LOST . For this, we first compute K-means on all the boxes produced by the first stage single object discovery model to obtain pseudo labels of the bounding boxes. Then a classical Faster RCNN is trained on the pseudo labels. As shown in Tab. 1, TokenCut with CAD outperforms the state-of-the-art by 5.7%, 4.9% and 5.1% on VOC07, VOC12 and COCO20k respectively.
In Fig. 3, we provide visualization for LOST , DSS and TokenCutMore visual results can be found in the project webpage.. For each method, we visualize the heatmap that is used to perform object detection. For LOST, the detection is mainly based on the map of inverse degree (). For DSS, the heatmap is the attention map associated to the second eigenvector. For TokenCut, we display the second smallest eigenvector. The visual results demonstrate that TokenCut can extract a high quality segmentation for the salient object. Compared with LOST and DSS, TokenCut is able to extract a more complete segmentation as can be seen in the first and the second samples in Fig. 3. In other cases, when LOST and DSS are unable to detect a large object, TokenCut can detect the object properly. Examples for this can be seen in the third and fourth samples in Fig. 3.
We further tested TokenCut on Internet imagesWe provide an online demo allowing to test Internet images.. The results are in Fig 5. It can be seen that even though the input images have noisy backgrounds, TokenCut can provide a precise attention map to cover the object and lead to an accurate prediction of the bounding box, demonstrates robustness of the method.
3 Unsupervised Saliency detection
We validated the performance of TokenCutfor unsupervised Saliency detection using three datasets : Extended Complex Scene Saliency Dataset(ECSSD) , DUTS and DUT-OMRON . ECSSD contains 1 000 real-world images of complex scenes for testing. DUTS contains 10 553 train and 5 019 test images. The training set is collected from the ImageNet detection train/val set. The test set is collected from ImageNet test, and the SUN dataset . Following the previous work , we report the performance on the DUTS-test subset. DUT-OMRON contains 5 168 images of high quality natural images for testing.
We report three standard metrics: F-measure, IoU and Accuracy. F-measure is a standard measure for saliency detection, computed as , where the Precision and Recall are defined using a binarized predicted mask and a ground truth mask. The is the maximum value of 255 uniformly distributed binarization thresholds. Following previous work , we set for consistency. IoU(Intersection over Union) score is computed based on the binary predicted mask and the ground-truth, the threshold is set to 0.5. Accuracy measures the proportion of pixels that have been correctly assigned to the object/background. The binarization threshold is set to 0.5 for masks.
Qualitative results are shown in Tab. 2. TokenCut significantly outperforms previous state-of-the-art methods. Adding BS or CRF refines the boundary of an object and further boosts the TokenCut performance, as can be seen in the visual results presented in Fig. 4.
4 Unsupervised Video Segmentation
We further evaluate TokenCut using three commonly used datasets for unsupervised video segmentation: DAVIS , FBMS and SegTV2 . DAVIS contains 50 high-resolution real-word videos, where 30 are for training and 20 are for validation. Pixel-wise annotations are depicted for the principle moving object within the scene for each frame. FBMS consists of 59 multiple moving object videos, providing 30 videos for testing with a total of 720 annotation frames. SegTV2 contains 14 full pixel-level annotated video for multiple objects segmentation. Following , we fuse the annotation of all moving objects into a single mask on FBMS and SegTV2 datasets for fair comparison.
We report performance using Jaccard index. The Jaccard index measures the intersection of union between an output segmentation M and the corresponding ground-truth mask G, which has been formulated as .
We compare TokenCut to the state-of-the art unsupervised video segmentation results in Tab. 3. TokenCut achieves competitive performances for this task. Note that DyStaB must be trained on the entire DAVIS training set and uses the pretrained model for evaluation with the FBMS and SegTV2 datasets. DeSprite learns an auto-encoder model to optimize on each individual video. In contrast, TokenCut does not require training and generalizes well for all three datasets. Visual results are illustrated in Fig. 6, TokenCut can precisely segment moving objects even in the case of challenging occlusions. Adding CRF as a post-processing further improves the boundary for segmented regionsThe segmentation results of entire videos can be found in the project webpage..
5 Analysis
In Tab. 4, we provide an analysis on defined in Eqn 4. The results indicate that the effects of variations in value are not significant and that a suitable threshold is = 0.2 for image input and = 0.3 for video input.
In Tab. 5, we provide an ablation study with different transformer backbones. The “-S” and “-B” are ViT small and ViT base architecture respectively. The “-16” and “-8” represents patch sizes 16 and 8 respectively. The “DeiT” is pre-trained supervised transformer model. The “MoCoV3” and “MAE” are pre-trained self-supervised transformer model. We optimise for different backbones: is set to 0.3 for MoCov3 and MAE, while for DINO and Deit is set to 0.2. Several insights can be found: 1) TokenCut is not suitable for supervised transformer models, while self-supervised transformers provide more powerful features allowing completing the task with TokenCut. 2) As LOST relies on a heuristic seeds expansion strategy, the performance varies significantly using different backbones. While our approach is more robust. Moreover, as no training is required for TokenCut, it might be a more straightforward evaluation for the self-supervised transformers.
In Tab. 6, we study different strategies to separate the nodes in into two groups using the second smallest eigenvector. We consider three natural methods: mean value (Mean), Expectation-Maximisation (EM), K-means clustering (K-means). We have also tried to search for the splitting point based on the best Ncut energy (Eqn 1). Note this approach is computationally expensive due to the quadratic complexity. The result suggests that the simple mean value as the splitting point performs well for most cases.
We also study the impact of using RGB or optical Flow for video segmentation. Quantitative results are presented in Tab. 7. We can see constructing graph on the entire video is better than constructing the graph per frame. We show an anaylsis by using the mean of optical flow feature and use it directly without feeding into transformer. The results illustrate that constructing the graph with RGB and RGB representation of the Flow together can significantly improve the performances over DAVIS . On FBMS and SegTV2 , due to the low quality of optical flow, the motion of salient objects are not detected in the optical flow as can be seen from the fact that they are not visible in the RGB visualisation. Some failure cases are shown in Fig. 8. This failure of optical flow to detect slow motion impedes the inference process for augmenting appearance with optical flow features. The low quality of optical flow can be attributed to three factors: 1) small motion between two frames; 2) low quality of raw image, for instance several examples in SegTV2, such as birdfall; 3) the absence of fine-tuning for the pre-trained optical flow model on these three datasets. Using both RGB appearance and flow lead to a slight improvement before edge refinement, but slightly worse results after edge refinement compared to using only RGB appearance. Some qualitative results are illustrated in Fig. 7. We can see how RGB frame and optical flow are complementary to each other: in the first row, the target moving person shares semantically similar features to other audiences and using only RGB frames would produce a mask cover all the persons; in the second row, the flow also has non-negligible values on the surface of the river, thus using only flow leads to worse performance.
In Tab. 8, we provide an analysis for different ways to construct graphs for video. For edges, we also consider the minimum and maximum values between the flow and RGB similarities. For nodes, a natural baseline is to build a graph for each single frame. We can see that the optimal choice is to use the average value of the flow and RGB similarities (Eqn. 4) and build a graph for an entire video.
Discussion
In the context of the Unsupervised Single Object Discovery task, the primary objective is to identify the most salient object within a given image. Consequently, we only choose the largest connected component in our approach. However, TokenCutcan identify more than one connected component in the second smallest eigenvector when multiple objects are present in the images. To illustrate this capability, we have included two examples in Fig 10. In Fig 11, we provide examples when multiple objects are moving from different directions. These results illustrate the robustness of our method.
Despite the good performance of the TokenCut proposal, it has several limitations. We show several failure cases in Fig. 9: i) As seen in the 1st row, TokenCut focuses on the largest salient part in the image, which may not be the desired object. ii) Similar to LOST , TokenCut assumes that a single salient object occupies the foreground. If multiple overlapping objects are present in an image, both LOST and our approach would fail to detect one of the object, as displayed in the 2nd row. iii) For object detection, neither LOST norTokenCut can handle occlusion properly, as shown in the 3rd row.
Conclusion
This paper describes TokenCut, an unified and effective approach for both image and video object segmentation without the need for supervised learning. TokenCut uses features from self-supervised transformers to constructs a graph where nodes are patches and edges represent similarities between patches. For videos, optical flow is incorporated to determine moving objects. We show that salient objects can be directly detected and delimited using the Normalized Cut algorithm. We evaluated this approach on unsupervised single object discovery, unsupervised saliency detection, and unsupervised video object segmentation, demonstrating that TokenCut can provide a significant improvement over previous approaches. Our results demonstrate that self-supervised transformers can provide a rich and general set of features that may likely be used for a variety of computer vision problems.
Acknowledgment
This work has been partially supported by the MIAI Multidisciplinary AI Institute at the Univ. Grenoble Alpes (MIAI@Grenoble Alpes - ANR-19-P3IA-0003), and by the EU H2020 ICT48 project Humane AI Net under contract EU #952026.