PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting
Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, Noah Snavely
Introduction
Vision as inverse graphics has long been an intriguing concept. Solving inverse rendering problems, i.e., recovering shape, material and lighting from images, has thus been a long-standing goal. Recently, neural rendering methods , have drawn significant attention due to their remarkable success in a range of problems, including shape reconstruction, novel view synthesis, non-physically-based relighting, and surface reflectance map estimation. These neural rendering methods adopt scene representations that are either physical, neural, or a mixture of both, along with a neural-network-based renderer. Methods that reconstruct textures or radiance fields work well for the task of interpolating novel views, but do not factorize appearance into lighting and materials, precluding physically-based appearance manipulation like material editing or relighting.
Prior multi-view inverse rendering methods assume RGBD input or varying illumination across input images achieved either by co-locating an active flashlight with moving cameras or capturing objects on a turntable with a fixed camera . Learning-based single-view methods that recover shape, illumination, and material properties have also been proposed .
In this work, we tackle the multi-view inverse rendering problem under the challenging setting of normal RGB input images sharing the same static illumination, without assuming scanned geometry. To this end, we propose PhySG, an end-to-end physically-based differentiable rendering pipeline to jointly estimate lighting, material, geometry and surface normals from posed multi-view images of specular objects. In our pipeline, we represent shape using signed distance functions (SDFs), building on their successful use in recent work . Additionally, a key component of our framework is our use of spherical Gaussians to approximate lighting and specular BRDFs allowing for efficient approximate evaluation of light transport . From 2D images alone, our method jointly reconstructs shape, illumination, and materials and allows for subsequent physics-based appearance manipulations such as material editing and relighting.
In summary, our contributions are as follows:
PhySG, an end-to-end inverse rendering approach to this problem of jointly estimating lighting, material properties, and geometry from multi-view images of glossy objects under static illumination. Our pipeline utilizes spherical Gaussians to approximately and efficiently evaluate the rendering equation in closed form.
Compared to prior neural rendering approaches, we show that PhySG not only generalizes to novel viewpoints, but also enables physically-intuitive material editing and relighting.
Background
Our approach lies at the intersection of multiple fields. We briefly review the related prior works below.
Neural rendering. The success of neural rendering has generated significant excitement. In particular, NeRF enables photo-realistic novel view synthesis by representing scenes as radiance fields via multi-layer-perceptrons (MLPs) and fitting these to a collection of input views. While NeRF represents scenes as volumetric opacity fields, other recent methods like DVR and IDR are surface-based. In these three works, appearance is represented by a single MLP that takes a 3D point (and a view direction), and outputs a color. Hence, their appearance model is essentially a surface light field that treats objects as light sources. Such an approach works well for novel view synthesis, but does not disentangle material and lighting, and hence is not suitable for physics-based relighting and material editing. Other approaches learn an appearance space from Internet photos of landmarks captured under diverse lighting, but are not physics-based and cannot generalize to arbitrary new lighting.
In contrast to such prior work that represents appearance as a single neural network, we model appearance via the physical rendering equation. Our approach can solve challenging inverse rendering problems involving specular or glossy objects under static lighting, and enable physically meaningful editing of lighting and materials.
Material and environment estimation. To estimate material properties, most prior works require scenes to be captured under varying illumination . They either place the object of interest on a mechanical turntable and capture it with a fixed camera , or move a camera with co-located flashlight to capture a static object from multiple viewpoints . The varying illumination yields rich cues for inferring material properties and geometry . For environment estimation from multi-view images, prior works factorize scene appearance into diffuse image and surface reflectance map given high-quality geometry from RGBD sensors. The surface reflectance map entangles the material and lighting, because it represents the distant environmental illumination convolved with an object’s specular BRDF, hence preventing relighting. In contrast to a global environment map, Azinovic et al. model lighting as surface emissions, and use a Monte Carlo differentiable renderer to jointly estimate material properties and surface emissions from multi-view images conditioned on scanned geometry and object segmentation masks. Other work seeks to predict illumination, materials and shape from a single image via learning-based priors . Ramamoorthi and Hanrahan estimates BRDF and lighting via deconvolution given known geometry. In our work, we aim to jointly estimate the material and environment, together with geometry and surface normals, solely from multi-view 2D images under the challenging setting of unknown static natural illumination.
Joint shape and appearance refinement. Given the initial geometry and appearance from RGBD sensors, Maier et al. and Zollhofer et al. jointly refine geometry and appearance by assuming Lambertian BRDF and incorporating shading cues. They pre-compute a lighting model based on spherical harmonics , then fix it while optimizing the shape and diffuse albedo. They adopt voxelized SDFs as their geometric representation. Assuming known illumination, Oxholm and Nishino also exploit reflectance cues to refine geometry computed via visual hulls. In contrast to these prior works, our method does not require scanned geometry or a known environment map. Instead, we estimate material and lighting parameters, as well as geometry and surface normals in an end-to-end fashion.
The rendering equation. Kajiya et al. proposed the rendering equation based on the physical law of energy conservation. For a surface point with surface normal , suppose is the incident light intensity at location along the direction , and BRDF is the reflectance coefficient of the material at location for incident light direction and viewing direction , then the observed light intensity is an integral over the hemisphere Viewing direction , lighting direction and surface normal are all assumed to point away from the scene.:
The BRDF is a function of viewing direction , and models view-dependent effects such as specularity.
Method
In this section, we describe our PhySG pipeline and its three major components: (1) geometry modeling, (2) appearance modeling, and (3) forward rendering. These components are designed to be differentiable, so that the whole pipeline can be optimized end-to-end from multiple images captured under static illumination.
Geometry modeling. Motivated by the success of signed distance functions (SDFs) for representing shape , we adopt SDFs as our geometric representation. SDFs support ray casting via sphere tracing, are differentiable, and automatically satisfy the constraint between shape and surface normal—the surface normal is exactly the gradient of the SDF. We represent SDFs with MLPs (rather than voxel grids) for their memory efficiency and infinite resolution . Concretely, let be our SDF, We assume SDF0 is an object’s exterior, while SDF0 is its interior. where is a 3D point and are the MLP weights. Our MLP consists of 8 nonlinear layers of width 512, with a skip connection at layer. To allow the MLP to model high-frequency geometric detail, we use positional encoding with 6 frequency components to encode the location of a 3D point . Using frequency components, positional encoding maps vector to . An alternate to SDFs is to use occupancy fields , but ray tracing through occupancy fields is much slower, requiring root-finding to locate the surface. While occupancy fields require over 100 MLP evaluations per cast ray , for SDFs, the MLP only needs to be evaluated 10 times via sphere tracing.
To render the pixel color for a camera ray, we first find the ray’s point of intersection with the SDF by starting from the ray’s intersection with the object bounding box and marching along the ray via sphere tracing as in , where the size of each step is the signed distance at the current location. The intersection point’s location and surface normal are then used by our appearance component to render the pixel’s color. Hence, to optimize the geometry, gradients must back-propagate through both and to the SDF parameters . Back-propagating through the surface normal is straightforward via auto-differentiation . To back-propagate through the surface location , we use the implicit differentiation method presented in . Note however that the sphere tracing algorithm itself need not be differentiable, hence it is very memory-efficient.
Appearance modeling. To model a single-material specular object in a way consistent with the rendering equation (see Eq. 1), we use two optimizable components: (1) an environment map, and (2) BRDF consisting of spatially varying diffuse albedo and a shared monochrome isotropic specular component. Note however that we do not model self-occlusion or indirect illumination. The hemispherical integral in the rendering equation generally does not have a closed-form expression, necessitating expensive Monte-Carlo methods for numeric evaluation. However, in our setting of glossy material and distant direct illumination, we can utilize spherical Gaussians (SGs) to efficiently approximate the rendering equation in closed form.
An -dimensional spherical Gaussian (SG) is a spherical function that takes the form :
We represent the spatially-varying diffuse albedo with an MLP mapping a surface point to a color vector , i.e., . Positional encoding is also applied to fit high-frequency texture details . Specifically, we use an MLP with 4 nonlinear layers of width 512, and encode location with frequencies. As for the shared specular component, we use the same simplified Disney BRDF model as in prior work :
where , accounts for the Fresnel and shadowing effects, and is the normalized distribution function. We include details of and in the supplemental material. We represent with a single SG:
Our isotropic specular BRDF assumption results in aligning with surface normal, i.e., , while the monochrome assumption makes the three numbers in identical.
To evaluate the rendering equation at a point with surface normal viewed along direction , must be spherically warped, while must be approximated by a constant at this specific location :
Hence for the point , we have:
Now that both and in the rendering equation are represented with SGs, we further approximate the remaining term with a SG :
Finally, we integrate the multiplication of these SGs in closed-form to compute the observed color .
To summarize, the optimizable parameters in our appearance component are , , and , which are parameters of the environment map, specular BRDF, and spatially-varying diffuse albedo, respectively.
Forward rendering. Given our geometric and appearance components, we perform forward rendering of a ray’s color as follows: (1) use sphere tracing to find the intersection point between the ray and the surface ; (2) compute the surface normal at via automatic differentiation ; (3) compute the diffuse albedo at ; (4) use the surface normal , environment map , diffuse albedo , specular BRDF , and viewing direction , to compute the color for ray by evaluating the rendering equation in closed form with our SG approximation. This procedure is illustrated in Fig. 2.
We now show that our pipeline is fully differentiable, in that its output (the rendered color) is differentiable w.r.t. all the optimizable parameters. First, the rendered color is differentiable w.r.t. the variables in step (4), because the SG renderer is simply the closed-form integration of spherical Gaussians. Since the diffuse albedo is an MLP in step (3), the rendered color is differentiable w.r.t. and by the chain rule. For our geometric model, we have shown that there exist gradients of both the surface location and surface normal w.r.t. the SDF parameters . Thus by the chain rule, the rendered color is differentiable w.r.t. as well.
where is a smooth approximation of a horizontally flipped ReLU (larger yields tighter approximation); and and are weights balancing different loss terms. We set in our experiments; gradually grows from 50 to 1600 as suggested in . Finally, rather than sampling independent pixels, we sample patches of size , and add an additional loss term to penalize the variance of surface normals inside patches consisting only of object pixels. We set the weight for this smoothness loss to 10. We train on a single 12GB NVIDIA GPU for 250k iterations.
Initialization. The SDF weights are initialized using the method of such that the initial shape is roughly a sphere. The diffuse albedo is initialized such that predicted albedo is 0.5 at all locations inside the object bounding box. For the specular BRDF, the initial lobe sharpness is randomly drawn from $\boldsymbol{\mu}[0.18,0.26]\sim$0.5. In addition, since different captures can vary significantly in exposure, we scale all input images of an object with the same constant such that the median intensity of all scaled images is 0.5. We empirically find that if the initial environment map is too bright or too dark, the diffuse albedo MLP sometimes gets stuck, predicting all zeros or ones during training. Our proposed initialization addresses this issue.
Experiments
We perform experiments on both synthetic and real-world data to validate our PhySG pipeline.
To create synthetic data, we use objects from ; for each object, we render 200 images with colored environmental lighting using the Mitsuba renderer , 100 each for training and testing. To test the extrapolation capability of different algorithms, the test images are distributed inside a 70-degree cone around the object’s north pole, while the training images cover the rest of the upper hemisphere (see Fig. 3). We use the Ward BRDF model included in Mitsuba, and set the specular albedo to and roughness values along the tangent and bitangent directions to . Ground truth surface normal maps and diffuse albedos are also rendered to quantitatively evaluate our inverse rendering results. To evaluate the relighting performance of our pipeline, we also render the same object with two other environment maps in Mitsuba to serve as ground truth.
We report image quality metrics: LPIPS , SSIM, and PSNR on held-out test viewpoints. As there is an inherent scale ambiguity in inverse rendering problems, before computing the metrics, we align our predicted image to ground-truth via channel-wise scale factors. Specifically, let denote the red channel of , respectively. Then the scale factor for the red channel is estimated via:
The green and blue channels are scaled similarly.
As shown by the quantitative evaluation in Tab. 1 and qualitative evaluation in Fig. 7, our synthesized novel test views, estimated diffuse albedo and surface normal, as well as material editing and relighting results closely match the ground truth on synthetic data, despite the test viewpoints representing a difficult view extrapolation scenario. Note especially that our method correctly extrapolates the challenging specular highlight in Fig. 7.
2 Real-world data
We test our method on multiple real-world captures from datasets including SLF , DeepVoxels , Bag of Chips and DTU . The objects in these captures are glossy and the illumination is static across different views.
SLF dataset. We use the glossy fish from . This dataset is captured with a gantry in a lab-controlled environment. The cameras are distributed on a hemisphere around the center object. We discard images in which the center object has noticeable shadows cast by the gantry or is partly occluded by the platform. Then we split the data according to the cameras’ latitudes, with test cameras’ latitudes above 55 degrees, as shown in Fig. 3. We render object segmentation masks from the provided laser-scanned meshes.
DeepVoxels. We use the glossy globe and coffee objects from DeepVoxels . These are real-world hand-held captures. The camera parameters are recovered with COLMAP . We use background removal tools to automatically generate the object segmentation masks. We leave 25% images for testing.
Bag of Chips. We use the glossy cans and corncho1 data from this dataset . We render object segmentation masks from the provided mesh scanned by RGBD sensors. We leave 25% images for testing.
DTU dataset. We use the shiny scan114 buddha object from this dataset . We discard images in which the camera casts noticeable shadows on the object. The object segmentation masks are automatically generated using background removal tools . We leave 25% images for testing.
Our inverse rendering results are qualitatively shown in Fig. 5. Video demos are shown in our supplemental material. We can see that our pipeline generates photo-realistic novel views, plausible material editing and relighting results.
3 Comparison with baselines
We could not identify prior work tackling exactly the same problem as us: simultaneously reconstructing lighting, material, and geometry from scratch from 2D images captured under static illumination. Hence, we compare our PhySG to the most related neural rendering approaches, including NeRF , IDR and DVR , in terms of novel view extrapolation quality.
Like PhySG, these approaches can also be trained end-to-end from 2D image supervision only, but they differ from our method in the way that appearance is modelled. Loosely speaking, they all model appearance as an MLP-represented surface light field. In particular, NeRF maps location and viewing direction to a color; IDR maps , , and surface normal to a color, and DVR only takes location . Tab. 3 and Fig. 4 compares different methods on synthetic and real data. NeRF does poorly in view extrapolation because its volumetric representation does not concentrate colors around surfaces as well as surface-based approaches. Although DVR uses a surface-based shape model, its model does not support view-dependent effects, and so it fails to model this kind of glossy data. Compared with DVR, IDR models view-dependence and does a good job in view extrapolation. However, it still has trouble synthesizing specular highlights, due to the lack of a physical model of appearance. In contrast, our method models such highlights well. As for geometry, our estimated geometry is nearly as good as IDR’s (and much better than other baselines) shown in Tab. 2, while we also allow for relighting and material editing.
We also tried redner , a Monte Carlo differentiable renderer representing shapes as meshes. We let redner jointly optimize lighting, texture, BRDF and geometry with an image reconstruction loss. We initialized the mesh a sphere at the beginning of training, and found that redner got stuck in the initial mesh and failed to converge.
4 Robustness to material roughness
Our method relies on specular highlights to estimate the lighting and material properties. As a result, if the material of interest is purely Lambertian, we face a lighting-texture ambiguity and cannot recover lighting without additional priors. We empirically test the robustness of our pipeline to material roughness on synthetic data. As shown in Fig. 6, even from very weak specular highlights, our method can reconstruct a reasonable-looking environment map.
Conclusion
We proposed PhySG, an end-the-end inverse rendering pipeline that uses physics-based differentiable rendering. PhySG uses signed distance functions (SDFs) and spherical Gaussians (SG) to represent geometry and appearance, respectively. We show that PhySG can jointly recover environment maps, material BRDFs and geometry from multi-view inputs captured under static illumination, enabling physics-based material editing and relighting.
Limitations. Our method has a few limitations that can be the subject of future work. First, indirect illumination is not modelled by our SG approximation of the rendering equation, which limits our method to object-level data. To lift this restriction and extend to scene-level data, differentiable path tracing combined with deferred neural textures can be explored, where only rough geometry is required for guidance. Second, we assume constant and monochrome specular BRDFs (with spatially-varying diffuse components). This assumption is due to the scale ambiguity between illumination and reflectance. Similar to intrinsic image decomposition, learning-based priors could help alleviate such ambiguities. Last, our work can also be extended to handle anisotropic or data-driven BRDFs, e.g., by fitting a mixture of anisotropic SGs .
Acknowledgements. This work was supported in part by the National Science Foundation (IIS-2008313, CHS-1900783, CHS-1930755).