Non-Local Spatial Propagation Network for Depth Completion

Jinsun Park, Kyungdon Joo, Zhe Hu, Chi-Kuei Liu, In So Kweon

Introduction

Depth estimation has become an important problem in recent years with the rapid growth of computer vision applications, such as augmented reality, unmanned aerial vehicle control, autonomous driving, and motion planning. To obtain a reliable depth prediction, information from various sensors is utilized, e.g., RGB cameras, radar, LiDAR, and ultrasonic sensors . Depth sensors, such as LiDAR sensors, produce accurate depth measurements with high frequency. However, the density of the acquired depth is often sparse due to hardware limitations, such as the number of scanning channels. To overcome such limitations, there have been a lot of works to estimate dense depth information based on the given sparse depth values, called depth completion.

Early methods for depth completion rely only on sparse measurement. Therefore, their predictions suffer from unwanted artifacts, such as blurry and mixed-depth values (i.e., mixed-depth problem). Because RGB images show subtle changes of color and texture, recent methods use RGB images as the guidance to predict accurate dense depth maps.

Direct depth completion algorithms take RGB or RGB-D images and directly infer a dense depth using a deep convolutional neural network (CNN). These direct algorithms have shown superior performance compared to conventional ones; however, they still generate blurry depth maps near depth boundaries. Soon after, this phenomenon is alleviated by recent affinity-based spatial propagation methods . By learning affinities for local neighbors and iteratively refining depth predictions, the final dense depth becomes more accurate. Nonetheless, previous propagation networks have an explicit limitation that they have a fixed-local neighborhood configuration for propagation. Fixed-local neighbors often have irrelevant information that should not be mixed with reference information, especially on depth boundaries. Hence, they still suffer from the mixed-depth problem in the depth completion task (see Fig. 1(c)).

To tackle the problem, we propose a Non-Local Spatial Propagation Network (NLSPN) that predicts non-local neighbors for each pixel (i.e., where the information should come from) and then aggregates relevant information using the spatially-varying affinities (i.e., how much information should be propagated), which are also predicted from the network. By relaxing the fixed-local neighborhood configuration, the proposed network can avoid irrelevant local neighbors affiliated with other adjacent objects. Therefore, our method is inherently robust to the mixed-depth problem. In addition, based on our analysis of conventional affinity normalization schemes, we propose a learnable affinity normalization method that has a larger representation capability of affinity combinations. It enables more accurate affinity estimation and thus improves the propagation among non-local neighbors. To further improve robustness to outliers from input and inaccurate initial prediction, we predict the confidence of the initial dense depth simultaneously, and it is incorporated into the affinity normalization to minimize the propagation of unreliable depth values. Experimental results on the indoor and outdoor datasets demonstrate that our method achieves superior depth completion performance compared with state-of-the-art methods.

Related Work

Depth Estimation and Completion The objective of depth estimation is to generate dense depth predictions based on various input information, such as a single RGB image, multi-view images, sparse LiDAR measurements, and so on. Conventional depth estimation algorithms often utilize information from a single modality. Eigen et al. used a multi-scale neural network to predict depth from a single image. In the method introduced by Zbontar and LeCun , the deep features of image patches are extracted from stereo rectified images, and then the disparity is determined by searching for the most similar patch along the epipolar line. Depth estimation with accurate but sparse depth information (i.e., depth completion) has been intensively explored as well. Uhrig et al. proposed sparsity invariant CNNs to predict a dense depth map given a sparse depth image from a LiDAR sensor. Ma and Sertac introduced a method to construct a 4D volume by concatenating RGB and sparse depth images and then feed it into an encoder-decoder CNN for the final prediction. Chen et al. adopted a fusion of 2D convolution and 3D continuous convolution to effectively consider the geometric configuration of 3D points.

Spatial Propagation Network Although direct depth completion algorithms have demonstrated decent performance, sparse-to-dense propagation with accurate guidance from different modalities (e.g., an RGB image) is a more effective way to obtain dense prediction from sparse inputs . Liu et al. proposed a spatial propagation network (SPN) to learn local affinities. The SPN learns task-specific affinity values from large-scale data, and it can be applied to a variety of high-level vision tasks, including depth completion and semantic segmentation. However, the individual three-way connection in four-direction is adopted for spatial propagation, which is not suitable for considering all local neighbors simultaneously. This limitation was overcome by Cheng et al. , who proposed a convolutional spatial propagation network (CSPN) to predict affinity values for local neighbors and update all the pixels simultaneously with their local context for efficiency. However, both the SPN and the CSPN rely on fixed-local neighbors, which could be from irrelevant objects. Therefore, the propagation based on those neighbors would result in mixed-depth values, and the iterative propagation procedure used in their architectures would increase the impact. Moreover, the fixed neighborhood patterns restrict the usage of relevant but wide-range (i.e., non-local) context within the image.

Non-Local Network The importance of non-local information has been widely explored in various vision tasks . Recently, a non-local block in deep neural networks was proposed by Wang et al. . It consists of pairwise affinity calculation and feature-processing modules. The authors demonstrated the effectiveness of non-local blocks by embedding them into existing deep networks for video classification and image recognition. These methods showed significant improvement over local methods.

Our Work Unlike previous algorithms , our network is trained to predict non-local neighbors with corresponding affinities. In addition, our learnable affinity normalization algorithm searches for the optimal affinity space, which has not been explored in conventional algorithms . Furthermore, we incorporate the confidence of the initial dense depth prediction (which will be refined by propagation procedure) into affinity normalization to minimize the propagation of unconfident depth values. Figure 2 shows an overview of our algorithm. Each component will be described in subsequent sections in detail.

Non-Local Spatial Propagation

The goal of spatial propagation is to estimate missing values and refine less confident values by propagating neighbor observations with corresponding affinities (i.e., similarities). Spatial propagation has been utilized as one of the key modules in various computer vision applications . In particular, spatial propagation is suitable for the depth completion task , and its superior performance compared to direct regression algorithms has been demonstrated . In this section, we first briefly review the local SPNs and their limitations, and then describe the proposed non-local SPN.

where (m,n)(m,n) and (i,j)(i,j) are the coordinates of reference and neighbor pixels, respectively; wm,ncw_{m,n}^{c} represents the affinity of the reference pixel; and wm,ni,jw_{m,n}^{i,j} indicates the affinity between the pixels at (m,n)(m,n) and (i,j)(i,j). The first term in the right-hand side represents the propagation of the reference pixel, while the second term stands for the propagation of its neighbors weighted by the corresponding affinities. The affinity of the reference pixel wm,ncw_{m,n}^{c} (i.e., how much the original value will be preserved) is obtained as

2 Non-Local Spatial Propagation Network

The SPN and the CSPN are effective in propagating information from more confident areas into less confident ones with data-dependent affinities. However, their potential improvement is inherently limited by the fixed-local neighborhood configuration (Fig. 3(e)). The fixed-local neighborhood configuration ignores object/depth distribution within the local area; thus, it often results in mixed-depth values of foreground and background objects after propagation. Although affinities predicted from the network can alleviate the depth mixing between irrelevant pixels to a certain degree, they can hardly avoid incorrect predictions and hold up the use of appropriate neighbors beyond the local area.

where I\mathbf{I} and D\mathbf{D} are the RGB and sparse depth images, respectively, and fϕ(⋅)f_{\phi}(\cdot) is the non-local neighbor prediction network that estimates KK neighbors for each pixel, under the learnable parameters ϕ\phi. We adopt an encoder-decoder CNN architecture for fϕ(⋅)f_{\phi}(\cdot), which will be described in Sec. 5.1. It should be noted that pp and qq are real numbers in Eq. (5); thus, the non-local neighbors can be defined to sub-pixel accuracy, as illustrated in Fig. 3(c).

Figure 3(f) shows some examples of appropriate and desired non-local neighbors near depth boundaries. In the fixed-local setup, affinity learning learns how to encourage the influence of the related pixels and suppress that of unrelated ones simultaneously. On the contrary, affinity learning with the non-local setup concentrates on relevant neighbors, and this facilitates the learning process.

Confidence-Incorporated Affinity Learning

Affinity learning is one of the key components in SPNs, which enables accurate and stable propagation. Conventional affinity-based algorithms utilize color statistics or hand-crafted features . Recent affinity learning methods adopt deep neural networks to predict affinities and show substantial performance improvement. In these methods, affinity normalization plays an important role to stabilize the propagation process.

In this section, we analyze the conventional normalization approach and its limitation, and then propose a normalization approach in a learnable way. Moreover, we incorporate the confidence of the initial prediction during normalization to suppress negative effects from unreliable depth values during propagation.

The purpose of affinity normalization is to ensure stability during propagation. For stability, the norm of the temporal Jacobian of xx, ∂xt/∂xt−1\partial x^{t}/\partial x^{t-1} should be equal to or less than one . Under the spatial propagation formulation in Eq. (1), this condition would be satisfied if ∑(i,j)∈Nm,n∣wm,ni,j∣≤1, ∀m,n\sum_{(i,j)\in\mathcal{N}_{m,n}}|w^{i,j}_{m,n}|\leq 1,\ \forall m,n. To enforce the condition, previous works normalize affinities by the absolute-sum (dubbed Abs−Sum\mathtt{Abs{-}Sum}) as follows:

where w^\hat{w} denotes the raw affinity before normalization. Although the stability condition is satisfied by Abs−Sum\mathtt{Abs{-}Sum}, it has a problem in that the viable combinations of normalized affinities are biased to a narrow high-dimensional space.

Without loss of generality, we first analyze the biased affinity problem using a toy example of the 2-neighbor case and then present solutions to the issue. In the 2-neighbor case, we denote affinities of the two neighbors as w1w_{1} and w2w_{2} with a slight abuse of notation. We assume that the unnormalized affinities are sampled from the standard normal distribution, N(0,1)\mathsf{N}(0,1) for simplicity.

For the Abs−Sum\mathtt{Abs{-}Sum}, the normalized affinities lie on the lines satisfying ∣w1∣+∣w2∣=1|w_{1}|+|w_{2}|=1 (referred to as A1A_{1}), as shown in Fig. 4(a). This limits the usage of potentially advantageous affinity configuration within the area ∣w1∣+∣w2∣<1|w_{1}|+|w_{2}|<1 (referred to as A2A_{2}). To fully explore the affinity configuration ∣w1∣+∣w2∣≤1|w_{1}|+|w_{2}|\leq 1, a simple remedy is to apply Eq. (6) only when ∑i∣wi∣>1\sum_{i}|w_{i}|>1 (noted as Abs−Sum∗\mathtt{Abs{-}Sum^{*}}). Figure 4(b) shows the affinity distribution of our simple remedy. However, the affinities normalized by Abs−Sum∗\mathtt{Abs{-}Sum^{*}} still have a high chance to fall on A1A_{1}. Indeed, with the increasing number of neighbors KK, the affinities are more likely to lie on A1A_{1}. (e.g., the normalization probability is 0.9850.985 when K=4K=4). Figure 4(e) (blue bars) shows the probability of affinities falling on A1A_{1} with various KK values.

One way to reduce the bias is to limit the range of raw affinities , for example, to [−1/C,1/C][-1/C,1/C] using the hyperbolic tangent function (tanh(⋅))(\texttt{tanh}(\cdot)) with a normalization factor CC. We refer to this normalization procedure as Tanh−C\mathtt{Tanh{-}C}, which is defined as follows:

where the condition C≥KC\geq K enforces the normalized affinities to guarantee ∑(i,j)∈Nm,n∣wm,ni,j∣≤1\sum_{(i,j)\in\mathcal{N}_{m,n}}|w^{i,j}_{m,n}|\leq 1; therefore, this condition ensures stability. Figure 4(c) shows the affinity distribution of Tanh−C\mathtt{Tanh{-}C} when C=2C=2. With a sacrifice of boundary values, Tanh−C\mathtt{Tanh{-}C} enables a more balanced affinity distribution. Moreover, the optimal value of CC in Tanh−C\mathtt{Tanh{-}C} may vary depending on the training task, e.g., the number of neighbors, the activation functions, and the dataset.

To determine the optimal value for the task, we propose to learn the normalization factor together with non-local affinities, and apply the normalization only when ∑(i,j)∈Nm,n∣wm,ni,j∣>1\sum_{(i,j)\in\mathcal{N}_{m,n}}|w^{i,j}_{m,n}|>1. The affinity of the proposed normalization, referred to as Tanh−γ−Abs−Sum∗\mathtt{Tanh{-}\gamma{-}Abs{-}Sum^{*}}, is defined as follows:

where γ\gamma denotes the learnable normalization parameter, and γmin\gamma_{min} and γmax\gamma_{max} are the minimum and maximum values that can be empirically set. Figure 4(d) shows an example of Tanh−γ−Abs−Sum∗\mathtt{Tanh{-}\gamma{-}Abs{-}Sum^{*}} when γ=1.25\gamma=1.25. Here, Tanh−γ−Abs−Sum∗\mathtt{Tanh{-}\gamma{-}Abs{-}Sum^{*}} can be viewed as a mixture of Abs−Sum∗\mathtt{Abs{-}Sum^{*}} and Tanh−C\mathtt{Tanh{-}C} (see Figs. 4(b) and (c)). The probability of affinities falling on the boundary with respect to the number of neighbors with γ=K/2\gamma=K/2 is shown in Fig. 4(e) (yellow bars). Compared to Abs−Sum∗\mathtt{Abs{-}Sum^{*}}, Tanh−γ−Abs−Sum∗\mathtt{Tanh{-}\gamma{-}Abs{-}Sum^{*}} still has a chance to avoid normalization, and it allows us to explore more diverse affinities with a larger number of neighbors.

2 Confidence-Incorporated Affinity Normalization

In the existing propagation frameworks , the affinity depicts the correlation between pixels and provides guidance for propagation based on similarity. In this case, each pixel in the map is treated equally without consideration of its reliability. However, in the depth completion task, different pixels should be weighted based on their reliability. For example, information from unreliable pixels (e.g., noisy pixels and pixels on depth boundaries) should not be propagated into neighbors regardless of their affinity to the neighboring pixels. The recent work DepthNormal addresses this problem with confidence prediction. It utilizes confidence as a mask for the weighted summation of input and prediction for seed point preservation. However, it does not fully prevent the propagation of incorrect depth values because weighted summation is conducted before each propagation separately.

In this work, we consider the confidence map of pixels and combine it with affinity normalization. That is, we predict not only the initial dense depth but also its confidence, and then the confidence is incorporated into affinity normalization to reduce disturbances from unreliable depths during propagation. The affinity of the confidence-incorporated Tanh−γ−Abs−Sum∗\mathtt{Tanh{-}\gamma{-}Abs{-}Sum^{*}} is defined as follows:

where ci,j∈c^{i,j}\in denotes the confidence of the pixel at (i,j)(i,j).

Figure 5(d) shows an example of a confidence-agnostic depth estimation result. Some noisy input depth points generate unreliable depth values with low confidences (see Fig. 5(c)). Without using confidence, the noisy and less confident pixels would harm their neighbor pixels during propagation and lead to unpleasing artifacts (see Fig. 5(d)). After the incorporation of confidence into normalization, our algorithm can successfully eliminate the impact of unconfident pixels and generate more accurate depth estimation, as shown in Fig. 5(e).

Depth Completion Network

In this section, we describe network architecture and loss functions for network training. The proposed NLSPN mainly consists of two parts: (1) an encoder-decoder architecture for the initial depth map, a confidence map and non-local neighbors prediction with their raw affinities, and (2) a non-local spatial propagation layer with a learnable affinity normalization.

The encoder-decoder part of the proposed network is built upon residual networks , and it extracts high-level features from RGB and sparse depth images. Additionally, we adopt the encoder-decoder feature connection strategy to simultaneously utilize low-level and high-level features.

In Fig. 2, we provide an overview of our algorithm. Features from the encoder-decoder network are shared for the initial dense depth, confidence, non-local neighbor, and raw affinity estimation. Then non-local spatial propagation is conducted in an iterative manner. As described in Sec. 3.2, non-local neighbors can have fractional coordinates. To better incorporate fractional coordinates into training, differentiable sampling is adopted during propagation. We note that our non-local propagation can be efficiently calculated by deformable convolutions . Therefore, each propagation requires a simple forward step of deformable convolution with our affinity normalization. Please refer to the supplementary material for the detailed network configuration.

2 Loss Function

Experimental Results

In this section, we first describe implementation details and the training environment. After that, quantitative and qualitative comparisons to previous algorithms on indoor and outdoor datasets are presented. We also present ablation studies to verify the effectiveness of each component of the proposed algorithm.

The proposed method was implemented using PyTorch with NVIDIA Apex and trained with a machine equipped with Intel Xeon E5-2620 and 4 NVIDIA GTX 1080 Ti GPUs. For all our experiments, we adopted an ADAM optimizer with β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, and the initial learning rate of 0.001. The network training took about 1 and 3 days on the NYU Depth V2 and KITTI Depth Completion datasets, respectively. We adopted the ResNet34 as our encoder-decoder baseline network. The number of non-local neighbors was set to 8 for a fair comparison to other algorithms using 3×33{\times}3 local neighbors. The number of propagation steps was set to 18 empirically. Other training details will be described for each dataset individually. For the quantitative evaluation, we utilized the following commonly used metrics :

RMSE (mm) : 1∣V∣∑v∈V∣ dvgt−dvpred ∣2\sqrt{\frac{1}{\left|\mathcal{V}\right|}\sum_{v\in\mathcal{V}}\left|\ d_{v}^{gt}-d_{v}^{pred}\ \right|^{2}}

MAE (mm) : 1∣V∣∑v∈V∣ dvgt−dvpred ∣\frac{1}{\left|\mathcal{V}\right|}\sum_{v\in\mathcal{V}}\left|\ d_{v}^{gt}-d_{v}^{pred}\ \right|

iRMSE (1/km) : 1∣V∣∑v∈V∣ 1/dvgt−1/dvpred ∣2\sqrt{\frac{1}{\left|\mathcal{V}\right|}\sum_{v\in\mathcal{V}}\left|\ 1/d_{v}^{gt}-1/d_{v}^{pred}\ \right|^{2}}

iMAE (1/km) : 1∣V∣∑v∈V∣ 1/dvgt−1/dvpred ∣\frac{1}{\left|\mathcal{V}\right|}\sum_{v\in\mathcal{V}}\left|\ 1/d_{v}^{gt}-1/d_{v}^{pred}\ \right|

REL : 1∣V∣∑v∈V∣ (dvgt−dvpred)/dvgt ∣\frac{1}{\left|\mathcal{V}\right|}\sum_{v\in\mathcal{V}}\left|\ (d_{v}^{gt}-d_{v}^{pred})/d_{v}^{gt}\ \right|

δτ\delta_{\tau} : Percentage of pixels satisfying max(dvgtdvpred,dvpreddvgt)<τ\texttt{max}\left(\frac{d_{v}^{gt}}{d_{v}^{pred}},\frac{d_{v}^{pred}}{d_{v}^{gt}}\right)<\tau

In Fig. 6, we present some depth completion results obtained for the NYUv2 dataset. As in previous works , 500 depth pixels were randomly sampled from a dense depth image and used as the input along with the corresponding RGB image. For comparison, we provide results from the Sparse-to-Dense (S2D) and the CSPN . The S2D (Fig. 6(c)) generates blurry depth images, as it is a direct regression algorithm. Compared to the S2D, the CSPN and our method generate depth maps with substantially improved accuracy thanks to the iterative spatial propagation procedure. However, the CSPN suffers from mixed-depth problems, especially on tiny or thin structures. In contrast, our method well preserves tiny structures and depth boundaries using non-local propagation.

Table 2 shows the quantitative evaluation of the NYUv2 dataset. The proposed algorithm achieves the best result and outperforms other methods by a large margin (RMSE 0.0200.020m). Compared to geometry-agnostic methods , geometry-aware ones show better performance in general. The proposed algorithm can be also viewed as a geometry-aware algorithm because it implicitly explores geometrically relevant neighbors for propagation.

2 KITTI Depth Completion

Table 2 shows the quantitative evaluation of the KITTI DC dataset. Similar to the results obtained for the NYUv2, geometry-aware algorithms perform better in general compared to geometry-agnostic methods . Since LiDAR sensor noise (i.e., mixed foreground and background points as shown in Fig. 5) is inevitable, the predicted confidence is highly beneficial to eliminate the impact of the noise. DepthNormal utilizes confidence values as a mask for weighted summation during refinement. However, its confidence mask does not totally prevent incorrect values from propagating into neighboring pixels. On the contrary, the proposed confidence-incorporated affinity normalization effectively restricts the propagation of erroneous values during propagation. We note that the proposed method outperformed all the peer-reviewed methods in the KITTI online leaderboard when we submitted the paper.

Figure 7 shows some examples of predicted dense depth with highlighted challenging areas. Those areas usually contain small structures near depth boundaries, which can be easily affected by the mixed-depth problem. Compared to the other methods (Figs. 7(c)-(g)), our algorithm (Fig. 7(h)) handles those challenging areas better with the help of non-local neighbors.

3 Ablation Studies

We conducted ablation studies to verify the role of each component of our network, including non-local propagation, affinity normalization, and the confidence-incorporated propagation. For all the experiments, we used a set of 10K images sampled from the KITTI DC training dataset for training and evaluated the performance on the full validation dataset. The network was trained for 20 epochs with center-cropped patches of 912×228912{\times}228 for fast training, and the batch size was set to 12. Other settings were set the same as those mentioned in Sec. 6.2.

Non-Local Neighbors Figure 3 visualizes some examples of non-local neighbors predicted by our algorithm. Compared to fixed-local neighbors, our predicted non-local neighbors have higher flexibility in the selection of neighbor pixels. In particular, non-local neighbors are selected from chromatically and geometrically relevant locations near the depth boundaries (e.g., same objects or planes). Moreover, we collected the statistics of the depth variance of neighboring pixels to show the relevance of the selected neighbors. On the KITTI DC validation set, the average depth variances for fixed-local and non-local neighbor configurations were 22.722.7mm and 11.611.6mm, respectively. The small variance of the non-local neighbor configuration demonstrates that the proposed method is able to select more relevant neighbors for propagation.

Affinity Normalization and Confidence Incorporation To validate the proposed affinity normalization algorithm, we compare it with three different affinity normalization methods (cf., Sec. 4). Table 3(g)-(i), and (m) assessed the performance using the same network but different affinity normalization methods. The model with Abs−Sum\mathtt{Abs{-}Sum} does not perform well due to the limited range of affinity combinations, as shown in Fig. 4(a). When relaxing the normalization condition while maintaining the stability condition (Abs−Sum∗\mathtt{Abs{-}Sum^{*}}), the performance was improved thanks to the wider area of feasible affinity space and better affinity distribution (Fig. 4(b)). Tanh−C\mathtt{Tanh{-}C} strengthens the stability condition without explicit normalization. However, as shown in Fig. 4(c), the resulting affinity values reside in a smaller affinity space (i.e., in a KK-dimensional hypercube with edge size 2/K2/K); therefore, it achieved a slightly worse performance compared to Abs−Sum∗\mathtt{Abs{-}Sum^{*}}. The proposed Tanh−γ−Abs−Sum∗\mathtt{Tanh{-}\gamma{-}Abs{-}Sum^{*}} was able to alleviate this limitation with a learnable normalization parameter γ\gamma. The learned γ\gamma compromises between Abs−Sum∗\mathtt{Abs{-}Sum^{*}} and Tanh−C\mathtt{Tanh{-}C}, and can boost the performance. Note that the final γ\gamma values (initialized with γ=K=8\gamma=K=8) trained on the NYUv2 (Sec. 6.1) and the KITTI DC (Sec. 6.2) datasets were 5.2 and 6.3, respectively. This observation indicates that the optimal γ\gamma varies based on the training environment.

Further Analysis To verify the importance of learned affinities, we further evaluated the proposed method with conventional affinities calculated based on the Euclidean distance between color intensities. As shown in Tab. 3(e) and (m), the network using learned affinities performed much better than the network using the hand-crafted one. In addition, we provide the number of network parameters of the compared methods in Tab. 4. The proposed method achieved superior performance with a relatively small number of network parameters. Please refer to the supplementary material for additional experimental results, visualizations, and ablation studies.

Conclusion

We have proposed an end-to-end trainable non-local spatial propagation network for depth completion. The proposed method gives high flexibility in selecting neighbors for propagation, which is beneficial for accurate propagation, and it eases the affinity learning problem. Unlike previous algorithms (i.e., fixed-local propagation), the proposed non-local spatial propagation efficiently excludes irrelevant neighbors and enforces the propagation to focus on a synergy between relevant ones. In addition, the proposed confidence-incorporated learnable affinity normalization encourages more affinity combinations and minimizes harmful effects from incorrect depth values during propagation. Our experimental results demonstrated the superiority of the proposed method.

Acknowledgement This work was partially supported by the National Information Society Agency for construction of training data for artificial intelligence (2100-2131-305-107-19).

References