Self-calibrating Deep Photometric Stereo Networks
Guanying Chen, Kai Han, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee K. Wong
Introduction
Photometric stereo aims at recovering the surface normal of a static object from a set of images captured under different light directions woodham1980ps; silver1980determining. Calibrated photometric stereo methods assume known light directions, and promising results have been reported shi2018benchmark at the cost of tedious light source calibration. The problem of uncalibrated photometric stereo, where light directions are unknown, still remains an open challenge, and its stable solution is wanted because of the ease of setting. In this work, we study the problem of uncalibrated photometric stereo for surfaces with general and unknown isotropic reflectance.
Most of the existing methods for uncalibrated photometric stereo alldrin2007r; shi2010self; papad14closed assume a simplified reflectance model, such as the Lambertian model, and focus on resolving the shape-light ambiguity, such as the Generalized Bas-Relief (GBR) ambiguity belhumeur1999bas. Although methods of lu2013uncalibrated; lu2015uncalibrated can handle surfaces with general bidirectional reflectance distribution functions (BRDFs), they rely on a uniform distribution of light directions for deriving a solution.
Recently, with the great success of deep learning in various computer vision tasks, deep learning based methods have been introduced to calibrated photometric stereo santo2017deep; Taniai18; ikehata2018cnn; chen2018ps. Instead of explicitly modeling complex surface reflectances, they directly learn the mapping from reflectance observations to surface normals given light directions. Although they have obtained promising results in a calibrated setting, they cannot handle the more challenging problem of uncalibrated photometric stereo, where light directions are unknown. One simple strategy to handle uncalibrated photometric stereo with deep learning is to directly learn the mapping from images to surface normals without taking the light directions as input. However, as reported in chen2018ps, the performance of such a model lags far behind those which take both images and light directions as input.
In this paper, we propose a two-stage model named Self-calibrating Deep Photometric Stereo Networks (SDPS-Net) to tackle this problem. The first stage of SDPS-Net, denoted as Lighting Calibration Network (LCNet), takes an arbitrary number of images as input and estimates their corresponding light directions and intensities. The second stage of SDPS-Net, denoted as Normal Estimation Network (NENet), estimates a surface normal map of a scene based on the lighting conditions estimated by LCNet and the input images. The rationales behind the design of our two-stage model are as follows. First, lighting information is very important for normal estimation since lighting is the source of various cues, such as shading and reflectance, and estimating the light directions (-vectors) and intensities (scalars) is in principle much easier than directly estimating the normal map (a -vector at each pixel location) together with the lighting conditions. Second, by explicitly learning to estimate light directions and intensities, the model can take advantage of the intermediate supervision, resulting in a more interpretable behavior. Last, the proposed LCNet can be seamlessly integrated with existing calibrated photometric stereo methods, which enables them to deal with unknown lighting conditions. Our code and model can be found at https://guanyingc.github.io/SDPS-Net.
Related Work
In this section, we review learning based photometric stereo and uncalibrated photometric stereo methods. We also briefly review the loosely related work on learning based lighting estimation. Readers are referred to shi2018benchmark for a comprehensive survey on calibrated photometric stereo with Lambertian surfaces and general BRDFs using non-learning based methods.
Recently, a few deep learning based methods have been introduced to calibrated photometric stereo santo2017deep; Taniai18; ikehata2018cnn; chen2018ps. Santo et al. santo2017deep proposed a fully-connected network to learn the mapping from reflectance observations captured under a pre-defined set of light directions to surface normal in a pixel-wise manner. Taniai and Maehara Taniai18 introduced an unsupervised learning framework that predicts both the surface normals and reflectance images of an object. Their model is “trained” at test time for each test object by minimizing the reconstruction loss between the input images and the rendered images. Ikehata ikehata2018cnn introduced a fixed shape representation, called observation map, that is invariant to the number and permutation of the images. For each surface point of the object, all its observations are merged into an observation map based on the given light directions, and the observation map is then fed to a convolutional neural network (CNN) to regress the normal vector. Chen et al. chen2018ps proposed a fully-convolutional network (FCN) to infer the normal map from the input image-lighting pairs, and an order-agnostic max-pooling operation was adopted to handle an arbitrary number of inputs. All the above methods assume known lighting conditions and cannot handle uncalibrated photometric stereo, where the light directions and intensities are not known a priori.
Uncalibrated photometric stereo
When lighting is unknown, the surface normals of a Lambertian object can only be estimated up to a linear ambiguity hayakawa1994photometric, which can be reduced to a -parameter GBR ambiguity belhumeur1999bas; yuille1999determining using the surface integrability constraint. Previous work used additional clues like albedo priors alldrin2007r; shi2010self, inter-reflections chandraker2005reflections, specular spikes drbohlav2005can, Torrance and Sparrow reflectance model georghiades2003incorporating, reflectance symmetry tan2007isotropy; wu2013calib, multi-view images esteban2008multiview, and local diffuse maxima papad14closed, to resolve the GBR ambiguity. Cho et al. cho2016photometric considered a semi-calibrated case where the light directions are known but not their intensities. There are few works that can handle non-Lambertian surfaces under unknown lighting. Hertzmann and Seitz hertzmann2005example proposed an exemplar based method by inserting an additional reference object to the scene. Methods based on cues like similarity in radiance changes sato2007shape; lu2013uncalibrated and attached shadow okabe2009attached were also introduced, but they require the light sources to be uniformly distributed on the whole sphere. Recently, Lu et al. lu2018symps introduced a method based on the “constrained half-vector symmetry” to work with non-uniform lightings. Different from these traditional methods, our method can deal with surfaces with general and unknown isotropic reflectance without the need of explicitly utilizing any additional clues or reference objects, solving a complex optimization problem at test time, or making assumptions on the light source distribution. The work most related to ours is the UPS-FCN introduced in chen2018ps. UPS-FCN is a single-stage model that directly regresses surface normals from images that are normalized by the known light intensities. Its performance lags far behind the calibrated methods. In contrast, our method solves the problem in two stages. We first tackles an easier problem of estimating the light directions and intensities, and then estimates the surface normals using the estimated lightings and the input images.
Learning based lighting estimation
Recently, learning based single-image lighting estimation methods have attracted considerable attention. Gardner et al. gardner2017learning introduced a CNN for estimating HDR environment lighting from an indoor scene image. Hold-Goeffroy et al. hold2017deep learned outdoor lighting using a physically-based sky model. Weber et al. weber2018learning estimated indoor environment lighting from an image of an object with known shape. Zhou et al. Zhou_2018_CVPR estimated lighting, in the form of Spherical Harmonics, from a human face image by assuming a Lambertian reflectance model. Different from the above methods, our method can estimate accurate directional lightings from multiple images of a static object with general shape and non-Lambertian surface.
Image Formation Model
where represents the measured intensity, accounts for attached shadows, and accounts for the global illumination effects (cast shadows and inter-reflections) and noise.
Based on this model, given the observations of surface points under different incoming lightings, the goal of uncalibrated photometric stereo is to estimate the surface normals for these surface points given only the measured intensities. In this work, we tackle this problem using a two-stage approach. In particular, we first estimate lightings from the measured intensities, and then solve for the surface normals using the estimated lightings and measured intensities.
Learning Uncalibrated Photometric Stereo
In this section, we introduce our two-stage framework, called SDPS-Net, for uncalibrated photometric stereo (see Fig. 1). The first stage of SDPS-Net, denoted as Lighting Calibration Network (LCNet, Fig. 1 (a)), takes an arbitrary number of images as input and estimates their corresponding light directions and intensities. The second stage of SDPS-Net, denoted as Normal Estimation Network (NENet, Fig. 1 (b)), estimates an accurate normal map of the object based on the lightings estimated by LCNet and the input images.
To estimate lightings from the images, an intuitive approach would be directly regressing the light direction vectors and intensity values. However, we propose that formulating the lighting estimation as a classification problem is a superior choice, as will be verified by our experiments. Our arguments are as follows. Fist, classifying a light direction into a certain range is easier than regressing the exact value(s), and this will reduce the learning difficulty. Second, taking discretized light directions as input may allow NENet to better tolerate small errors in the estimated light directions.
Since we cast our lighting estimation as a classification problem, we need to discretize the continuous lighting space. Note that a light direction in the upper-hemisphere can be described by its azimuth and elevation (see Fig. 2 (a)). We can discretize the light directoin space by evenly dividing both the azimuth and elevation into bins, resulting in classes (see Fig. 2 (b)). Solving a -class classification problem is not computationally efficient, as the softmax probability vector will have a very high dimension even when is not large (e.g., when ). Instead, we estimate the azimuth and elevation of a light direction separately, leading to two -class classification problems. Similarly, we evenly divide the range of possible light intensities into classes (e.g., for a possible light intensity range of ).
Local-global feature fusion
A straightforward approach to estimate the lighting for each image is simply taking a single image as input, encoding it into a feature map using a CNN, and feeding the feature map to a lighting prediction layer. It is not surprising that the result of such a simple solution is far from satisfactory. Note that the appearance of an object is determined by its surface geometry, reflectance model and the lighting. The feature map extracted from a single observation obviously does not provide sufficient information for resolving the shape-light ambiguity. Thanks to the nature of photometric stereo where multiple observations of an object are considered, we propose a local-global feature fusion strategy to extract more comprehensive information from multiple observations.
Specifically, we separately feed each image into a shared-weight feature extractor to extract a feature map, which we call local feature as it only provides information from a single observation. All local features of the input images are then aggregated into a global feature through a max-pooling operation, which has been proven to be efficient and robust on aggregating salient features from a varying number of unordered inputs wiles2017silnet; chen2018ps. Such a global feature is expected to convey implicit surface geometry and reflectance information of the object which help resolve the ambiguity in lighting estimation. Each local feature is concatenated with the global feature, and fed to a shared-weight lighting estimation sub-network to predict the lighting for each individual image. By taking both local and global features into account, our model can produce much more reliable results than using the local features alone. We empirically found that additionally including the object mask as input can effectively improve the performance of lighting estimation, as will be seen in the experiment section.
Network architecture
LCNet is a multi-input-multi-output (MIMO) network that consists of a shared-weight feature extractor, an aggregation layer (i.e., max-pooling layer), and a shared-weight lighting estimation sub-network (see Fig. 1 (a)). It takes the observations of the object together with the object mask as input, and outputs the light directions and intensities in the form of softmax probability vectors of dimension (azimuth), (elevation) and (intensity), respectively. We convert the output of LCNet to -vector light directions and scalar intensity values by simply taking the middle value of the range with the highest probability We have experimentally verified that alternative ways like taking the expectation of the probability vector or performing quadratic interpolation in the neighborhood of the peak value do not improve the result..
Loss function
Multi-class cross entropy loss is adopted for both light direction and intensity estimation, and the overall loss function is
where and are the loss terms for azimuth and elevation of the light direction, and is the loss term for light intensity. During training, weights , and for the loss terms are set to .
2 Normal Estimation Network
NENet is a multi-input-single-output (MISO) network. The network architecture of NENet is similar to PS-FCN chen2018ps, consisting of a shared-weight feature extractor, an aggregation layer, and a normal regression sub-network (see Fig. 1 (b)). The key difference between NENet and PS-FCN is that PS-FCN requires accurate lightings as input, whereas NENet is trained with discretized lightings estimated by the LCNet and shows a more robust behavior over noise in the lightings.
NENet first normalizes the input images using the light intensities predicted by LCNet, and then concatenates the light directions predicted by LCNet with the images to form the input of the shared-weight feature extractor. Given an image of size , the loss function for NENet is
3 Training Data
We adopted the publicly available synthetic Blobby and Sculpture datasets chen2018ps for training. Blobby and Sculpture datasets provide surfaces with complex normal distributions and diverse materials from MERL dataset matusik2003merl. Effects of cast shadow and inter-reflection were considered during rendering using the physically based raytracer Mitsuba jakob2010mitsuba. There are samples in total. Each sample was rendered under distinct light directions sampled from the upper-hemisphere with uniform light intensity, resulting in images (). The rendered images have a dimension of .
To simulate images under different light intensities, we randomly generated light intensities in the range of to scale the magnitude of the images (i.e., the ratio of the highest light intensity to the lowest one is ) Note that the ratio (other than the exact value) matters, since light intensity can only be estimated up to a scale factor.. Note that this selected range contains a wider range of intensity value than the public photometric stereo datasets like DiLiGenT benchmark shi2018benchmark and Gourd&Apple dataset alldrin2008p. The color intensities of the input images were normalized to the range of $[-0.025,0.025]128\times 128128\times 128$ as it contains fully-connected layers and requires the input to have a fixed spatial dimension. Trained only on the synthetic dataset, we will show that our model can generalize well on real datasets.
Experimental Results
We performed network analysis for our method, and compared our method with the previous state-of-the-art methods on both synthetic and real datasets.
Our framework was implemented in PyTorch paszke2017pytorch and Adam optimizer kingma2014adam was used with default parameters. LCNet and NENet contain million and million parameters, respectively. We first trained LCNet using a batch size of for epochs until convergence, and then trained NENet from scratch given the lightings estimated by LCNet with a batch size of for epochs. We found that end-to-end fine-tuning did not improve the performance. The learning rate was initially set to and halved every and epochs for LCNet and NENet, respectively. It took about hours to train LCNet and hours to train NENet on a single Titan X Pascal GPU with a fixed input image number of .
Evaluation metrics
To measure the accuracy of the predicted light directions and surface normals, the widely used mean angular error (MAE) in degree is adopted. Since the light intensities among the testing images can only be estimated up to a scale factor , we introduce the scale-invariant relative error
1 Network Analysis with Synthetic Data
To quantitatively perform network analysis for our method, we rendered a synthetic dataset, denoted as MERL, of sphere and bunny shapes, denoted as Sphere and Bunny hereafter, respectively, using the physically based raytracer Mitsuba jakob2010mitsuba. Each shape was rendered with isotropic BRDFs from MERL dataset matusik2003merl under light directions sampled from the upper-hemisphere, leading to test objects (see Fig. 3). Cast shadows and inter-reflections are considered for Bunny. For all experiments on synthetic dataset involving input with unknown light intensities, we randomly generated light intensities in the range of . Each experiment was repeated five times and the average results were reported.
Discretization of lighting space
For a given number of bins , the maximum deviation angle for azimuth and elevation of a light direction is after discretization (e.g., when ). To investigate how light direction discretization affects the normal estimation accuracy, we adopted the state-of-the-art calibrated method PS-FCN chen2018ps and MERL dataset as the testbed. We divided the azimuth and elevation of light directions into different number of bins ranging from to . For a specific bin number, we replaced each ground-truth light direction by each of the four light directions having the maximum possible angular deviations after discretization (see Fig. 4 (a)), respectively. We then used those light directions as input for PS-FCN to infer surface normals. The normal estimation error reported in Fig. 4 (b) is the upper-bound error for PS-FCN caused by the discretization. We can see that the increase in error caused by discretization is marginal when . In our implementation, we empirically set and to and , respectively. We experimentally found that the performance of LCNet is robust to different discretization levels. We chose a relatively sparse discretization of lighting space in this paper as it may allow NENet to learn to better tolerate small errors in the estimated lighting at test time.
Effectiveness of LCNet
To validate the design of LCNet, we compared LCNet with three baseline models for lighting estimation. The first baseline model, denoted as LCNet, is a regression based model that directly regresses the light direction vectors and intensity values (please refer to the supplementary for implementation details). The second baseline model, denoted as LCNet, is a classification based model that only takes the images as input without the object mask input. The last baseline model, denoted as LCNet, is a classification based model that independently estimates lighting for each observation (i.e., without local-global feature fusion). All models were trained under the same setting, and the results are summarized in Table 1.
Experiments with IDs A0 & A1 in Table 1 show that the proposed classification based LCNet consistently outperformed the regression based baseline on both light direction and intensity estimation. This echoes our hypothesis that classifying a light direction to a certain range is easier than regressing an exact value. Thus, solving the classification problem reduces the learning difficulty and improves the performance. Experiments with IDs A0 & A2 show that taking the object mask as input can effectively improve the lighting estimation results. This might be explained by the fact that object mask provides strong information for occluding contours of the object, and helps the network distinguish the shadow region from the non-object region. Experiments with IDs A0 & A3 show that the proposed local-global feature fusion strategy can effectively make use of information from multiple observations, and significantly improve the lighting estimation accuracy. Please refer to our supplementary for detailed lighting estimation results of LCNet on Bunny from MERL dataset.
Effectiveness of NENet
Experiments with IDs B1 & B2 in Table 2 show that after training with the discretized lightings estimated by LCNet, NENet performs better than PS-FCN given possibly noisy lightings at test time, while experiments with IDs B3 & B4 show that training NENet with the light directions estimated by the regression based baseline is not always helpful. This result further demonstrates that the proposed framework is robust to noisy lightings. Experiments with IDs B0 & B1 show the proposed method achieved results comparable to the fully calibrated method PS-FCN chen2018ps, with average MAEs of and on Sphere and Bunny, respectively.
Figure 5 shows that the performances of LCNet and NENet increased with the number of input images. This is expected, since more useful information can be used to infer the lightings and normals with more input images.
Comparison with single-stage models
To validate the effectiveness of the proposed two-stage framework, we compared our method with five different single-stage baseline models. We first retrained UPS-FCN chen2018ps, denoted as UPS-FCN, with images scaled by randomly generated light intensities to allow it adapt to unknown intensities at test time. We then increased the model capacity of UPS-FCN by introducing a wider network (i.e., more channels in the convolutional layers) and a deeper network (i.e., more convolutional layers), denoted as UPS-FCN and UPS-FCN, respectively. We also trained a deeper network, denoted as UPS-FCN, that takes both the images and object mask as input. We last investigated the effect of having additional lighting supervision by training a variant model, denoted as UPS-FCN, to simultaneously estimate lighting and surface normal. Please refer to our supplementary for detailed network architectures.
Experiments with IDs B5-B9 in Table 2 show that utilizing a wider or deeper network, taking the object mask as input, or incorporating additional lighting supervision can improve the performance of single-stage model in some extent. However, experiments with IDs B1 & B5 show that the proposed method significantly outperformed the best-performing single-stage model, especially on surfaces with complex geometry such as Bunny, when the input as well as the number of parameters are comparable. This result indicates that simply increasing the layer numbers or channel numbers of the network, or incorporating additional lighting supervision cannot produce optimal results.
Comparison with the non-learning method papad14closed
To further verify the effectiveness of our method over non-learning method, we compared SDPS-Net with the existing uncalibrated method PF14 papad14closed, which achieved state-of-the-art results on the DiLiGenT benchmark shi2018benchmark, on different lighting distributions and types of BRDFs. Specifically, we considered one near uniform and one biased lighting distribution (see Fig. 6 (a)). We rendered Bunny using four typical types of BRDFs, including the Lambertian model and three other types from MERL dataset matusik2003merl, namely, Fabric, Plastic, and Phenolic. They contained , , , and different BRDFs, respectively. We reported the average results for each type (see Fig. 6 (b) for an example of each type.).
Figures 6 (c)-(e) compare SDPS-Net and PF14 on lighting estimation and normal estimation. The following observations are made: 1) PF14 performed well on light direction and normal estimation for diffuse or near diffuse surfaces (i.e., Lambertian and Fabric), but will quickly degenerate when dealing with non-Lambertian surfaces. Besides, it cannot reliably estimate light intensities for all the BRDFs. 2) SDPS-Net performed well on different types of BRDFs, especially on surfaces exhibit specular highlights. This result suggests that specular highlight is an important clue for uncalibrated photometric stereo drbohlav2005can. 3) The performance of light direction and normal estimation of both methods will have a trend of decreasing when dealing with biased lighting distribution, while the performance of intensity estimation will slightly improve.
2 Evaluation on Real Datasets
We evaluated our method on three publicly available non-Lambertian photometric stereo datasets, namely the DiLiGenT benchmark shi2018benchmark, Gourd&Apple dataset alldrin2008p and Light Stage Data Gallery einarsson2006relighting. Figure 7 visualizes the lighting distribution of these datasets (note that for Light Stage Data Gallery, we only used images with the front side of the object under illumination). Since Gourd&Apple dataset and Light Stage Data Gallery only provide calibrated lightings (without ground-truth normal maps), we quantitatively evaluated our method on lighting estimation while qualitatively evaluated it on normal estimation.
Evaluation on DiLiGenT benchmark
Table 3 (a)-(b) show that LCNet outperformed the regression based baseline LCNet and achieved highly accurate results on both light direction and intensity estimation on DiLiGenT benchmark, with an average MAE of and an average relative error of , respectively. Table 3 (c) compares the normal estimation results of SDPS-Net with previous state-of-the-art methods on DiLiGenT benchmark. SDPS-Net achieved state-of-the-art results on almost all objects with an average MAE of , except for the Bear object. Although UPS-FCN achieved reasonably good results on objects with smooth surface and uniform material (e.g., Ball), it had difficulties in handling surfaces with complex geometry and spatially-varying BRDFs (e.g., Reading and Harvest). The normal estimation network coupled with LCNet (i.e., SDPS-Net) outperforms that with LCNet (i.e., LCNet+NENet†) with a clear improvement of in average MAE, demonstrating the effectiveness of the proposed classification based LCNet. It is interesting to see that, coupled with our LCNet, the calibrated methods L2 baseline woodham1980ps and IS18 ikehata2018cnn can already achieve results comparable to the previous state-of-the-art methods. This result indicates that our proposed LCNet can be integrated with existing calibrated methods to help handle cases where lighting conditions are unknown. Figures 8 (a)-(b) show the qualitative results of SDPS-Net on DiLiGenT benchmark.
Evaluation on other real datasets
Table 8(e) shows that SDPS-Net can estimate accurate light directions and intensities for the challenging Gourd&Apple dataset and Light Stage Data Gallery. Our method can also reliably recover visually pleasing surface normal of these two datasets (see Fig. 8 (c)-(g)), clearly demonstrating the practicality of the proposed methods in real world applications. Please refer to our supplementary for more results.
Conclusion and Discussion
In this paper, we have proposed a two-stage deep learning framework, called SDPS-Net, for uncalibrated photometric stereo. The first stage of our framework takes an arbitrary number of images as input and estimates their corresponding light directions and intensities, while the second stage predicts the normal map of the object based on the lightings estimated in the first stage and the input images. By explicitly learning to estimate lighting conditions, our two-stage framework can take advantage of the intermediate supervision to reduce the learning difficulty and improve the final normal estimation results. Besides, the first stage of our framework can be seamlessly integrated with existing calibrated methods, which enables them to handle uncalibrated photometric stereo. Experiments on both synthetic and real datasets showed that our method significantly outperformed existing state-of-the-art uncalibrated photometric stereo methods.
Since our framework is trained only on surfaces with uniform material, it may not perform well in dealing with steep color changes caused by multi-material surfaces (see Fig. 8 (b) for an example). In the future, we will investigate better training datasets and network architectures for handling surfaces with spatially-varying BRDFs.
We gratefully acknowledge the support of NVIDIA Corporation with the donation of the Titan X Pascal GPU. Kai Han is supported by EPSRC Programme Grant Seebibyte EP/M013774/1. Boxin Shi is supported in part by National Science Foundation of China under Grant No. 61872012. Yasuyuki Matsushita is supported by the New Energy and Industrial Technology Development Organization (NEDO).