Editing Conditional Radiance Fields

Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, Bryan Russell

Introduction

3D content creation often involves manipulating high-quality 3D assets for visual effects or augmented reality applications, and part of a 3D artist’s workflow consists of making local adjustments to a 3D scene’s appearance and shape . Explicit representations give artists control of the different elements of a 3D scene. For example, the artist may use mesh processing tools to make local adjustments to the scene geometry or change the surface appearance by manipulating a texture atlas . In an artist’s workflow, such explicit representations are often created by hand or procedurally generated.

While explicit representations are powerful, there remain significant technical challenges in automatically acquiring a high-quality explicit representation of a real-world scene due to view-dependent appearance, complex scene topology, and varying surface opacity. Recently, implicit continuous volumetric representations have shown high-fidelity capture and rendering of a variety of 3D scenes and overcome many of the aforementioned technical challenges . Such representations encode the captured scene in the weights of a neural network. The neural network learns to render view-dependent colors from point samples along cast rays, with the final rendering obtained via alpha compositing . This representation enables many photorealistic view synthesis applications . However, we lack critical knowledge in how to enable artists’ control and editing in this representation.

Editing an implicit continuous volumetric representation is challenging. First, how can we effectively propagate sparse 2D user edits to fill the entire corresponding 3D region in this representation? Second, the neural network for an implicit representation has millions of parameters. It is unclear which parameters control the different aspects of the rendered shape and how to change the parameters according to the sparse local user input. While prior work for 3D editing primarily focuses on editing an explicit representation , they do not apply to neural representations.

In this paper, we study how to enable users to edit and control an implicit continuous volumetric representation of a 3D object. As shown in Figure 1, we consider three types of user edits: (i) changing the appearance of a local part to a new target color (e.g., changing the chair seat’s color from beige to red), (ii) modifying the local shape (e.g., removing a chair’s wheel or swapping in new arms from a different chair), and (iii) transferring the color or shape from a target object instance. The user performs 2D local edits by scribbling over the desired location of where the edit should take place and selecting a target color or local shape.

We address the challenges in editing an implicit continuous representation by investigating how to effectively update a conditional radiance field to align with a target local user edit. We make the following contributions. First, we learn a conditional radiance field over an entire object class to model a rich prior of plausible-looking objects. Unexpectedly, this prior often allows the propagation of sparse user scribble edits to fill a selected region. We demonstrate complex edits without the need to impose explicit spatial or boundary constraints. Moreover, the edits appear consistently when the object is rendered from different viewpoints. Second, to more accurately reconstruct shape instances, we introduce a shape branch in the conditional radiance field that is shared across object instances, which implicitly biases the network to encode a shared representation whenever possible. Third, we investigate which parts of the conditional radiance field’s network affect different editing tasks. We show that shape and color edits can effectively take place in the later layers of the network. This finding motivates us to only update these layers and enables us to produce effective user edits with significant computational speed-up. Finally, we introduce color and shape editing losses to satisfy the user-specified targets, while preserving the original object structure.

We demonstrate results on three shape datasets with varying levels of appearance, shape, and training view complexity. We show the effectiveness of our approach for object view synthesis as well as color and shape editing, compared to prior neural editing methods. Moreover, we show that we can edit the appearance and shape of a real photograph and that the edit propagates to extrapolated novel views. We highly encourage viewing our video to see our editing demo in action. Code and more results are available at our GitHub repo and website.

Related Work

Our work is related to novel view synthesis and interactive appearance and shape editing, which we review here.

Novel view synthesis. Photorealistic view synthesis has a storied history in computer graphics and computer vision, which we briefly summarize here. The goal is to infer the scene structure and view-dependent appearance given a set of input views. Prior work reasons over an explicit or discrete volumetric representation of the underlying geometry. However, both have fundamental limitations – explicit representations often require fixing the structure’s topology and have poor local optima, while discrete volumetric approaches scale poorly to higher resolutions.

Instead, several recent approaches implicitly encode a continuous volumetric representation of shape or both shape and view-dependent appearance in the weights of a neural network. These latter approaches overcome the aforementioned limitations and have resulted in impressive novel-view renderings of complex real-world scenes. Closest to our approach is Schwarz et al. , where they build a generative radiance field over an object class and include latent vectors for the shape and appearance of an instance. Different from their methods, we include an instance-agnostic branch in our neural network, which inductively biases the network to capture common features across the shape class. As we will demonstrate, this inductive bias more accurately captures the shape and appearance of the class. Moreover, we do not require an adversarial loss to train our network and instead optimize a photometric loss, which allows our approach to directly align to a single view of a novel instance. Finally, our work is the first to address the question of how to enable a user to make local edits in this new representation.

Interactive appearance and shape editing. There has been much work on interactive tools for selecting and cloning regions and editing single still images . Recent works have focused on integrating user interactions into deep networks either through optimization or a feed-forward network with user-guided inputs . Here, we are concerned with editing 3D scenes, which has received much attention in the computer graphics community. Example interfaces include 3D shape drawing and shape editing using inflation heuristics , stroke alignment to a depicted shape , and learned volumetric prediction from multi-view user strokes . There has also been work to edit the appearance of a 3D scene, e.g., via transferring multi-channel edits to other views , scribble-based material transfer , editing 3D shapes in a voxel representation , and relighting a scene with a paint brush interface . Finally, there has been work on editing light fields . We encourage the interested reader to review this survey on artistic editing of appearance, lighting, and material . These prior works operate over light fields or explicit/discrete volumetric geometry whereas we seek to incorporate user edits in learned implicit continuous volumetric representations.

A closely related concept is edit propagation , which propagates sparse user edits on a single image to an entire photo collection or video. In our work, we aim to propagate user edits to volumetric data for rendering under different viewpoints. Also relevant is recent work on applying local “rule-based” edits to a trained generative model for images . We are inspired by the above approaches and adapt it to our new 3D neural editing setting.

Editing a Conditional Radiance Field

Our goal is to allow user edits of a continuous volumetric representation of a 3D scene. In this section, we first describe a new neural network architecture that more accurately captures the shape and appearance of an object class. We then describe how we update network weights to achieve color and shape editing effects.

To achieve this goal, we build upon the recent neural radiance field (NeRF) representation . While the NeRF representation can render novel views of a particular scene, we seek to enable editing over an entire shape class, e.g., “chairs”. For this, we learn a conditional radiance field model that extends the NeRF representation with latent vectors over shape and appearance. The representation is trained over a set of shapes belonging to a class, and each shape instance is represented by latent shape and appearance vectors. The disentanglement of shape and appearance allows us to modify certain parts of the network during editing.

Let x=(x,y,z){\bm{x}}=(x,y,z) be a 3D location, d=(ϕ,θ){\bm{d}}=(\phi,\theta) be a viewing direction, and z(s){\bm{z}}^{(s)} and z(c){\bm{z}}^{(c)} be the latent shape and color vectors, respectively. Let (c,σ)=F(x,d,z(s),z(c))\left({\bm{c}},\sigma\right)=\mathcal{F}{\left({\bm{x}},{\bm{d}},{\bm{z}}^{(s)},{\bm{z}}^{(c)}\right)} be the neural network for a conditional radiance field that returns a radiance c=(r,g,b){\bm{c}}=(r,g,b) and a scalar density σ\sigma. The network F\mathcal{F} is parametrized as a multi-layer perceptron (MLP) such that the density output σ\sigma is independent of the viewing direction, while the radiance c{\bm{c}} depends on both position and viewing direction.

To obtain the color at a pixel location for a desired camera location, first, NcN_{c} 3D points {ti}i=1Nc\{t_{i}\}_{i=1}^{N_{c}} are sampled along a cast ray r{\bm{r}} originating from the pixel location (ordered from near to far). Next, the radiance and density values are computed at each sampled point with network F\mathcal{F}. Finally, the color is computed by the “over” compositing operation . Let αi=1−exp⁡(−σiδi)\alpha_{i}=1-\exp{\left(-\sigma_{i}\delta_{i}\right)} be the alpha compositing value of sampled point tit_{i} and δi=ti+1−ti\delta_{i}=t_{i+1}-t_{i} be the distance between the adjacent sampled points. The compositing operation, which outputs pixel color C^\hat{C}, is the weighted sum:

Next, we describe details of our network architecture and our training and editing procedures.

NeRF finds the inductive biases provided by positional encodings and stage-wise network design critical. Similarly, we find the architectural design choices important and aim for a modular model, providing an inductive bias for shape and color disentanglement. These design choices allow for selected submodules to be finetuned during user editing (discussed further in the next section), enabling more efficient downstream editing. We illustrate our network architecture F\mathcal{F} in Figure 2.

First, we learn a category-specific geometric representation with a shared shape network Fshare\mathcal{F}_{\text{share}} that only operates on the input positional encoding γ(x)\gamma({\bm{x}}) . To modify the representation for a specific shape, an instance-specific shape network Finst\mathcal{F}_{\text{inst}} is conditioned on both the shape code z(s){\bm{z}}^{(s)} and input positional encoding. The representations are added and modified by a fusion shape network Ffuse\mathcal{F}_{\text{fuse}}. To obtain the density prediction σ\sigma, the output of Ffuse\mathcal{F}_{\text{fuse}} is passed to a linear layer, the output density network Fdens\mathcal{F}_{\text{dens}}. To obtain the radiance prediction c{\bm{c}}, the output of Ffuse\mathcal{F}_{\text{fuse}} is concatenated with the color code z(c){\bm{z}}^{(c)} and encoded viewing direction γ(d)\gamma({\bm{d}}) and passed through a two-layer MLP, the output radiance network Frad\mathcal{F}_{\text{rad}}. We follow Mildenhall et al. for training and jointly optimize the latent codes via backpropagation through the network. We provide additional training details in the appendix.

2 Editing via Modular Network Updates

We are interested in editing an instance encoded by our conditional radiance field. Given a rendering by the network F\mathcal{F} with shape zk(s){\bm{z}}^{(s)}_{k} and color zk(c){\bm{z}}^{(c)}_{k} codes, we desire to modify the instance given a set of user-edited rays. We wish to optimize a loss Ledit(F,zk(s),zk(c))\mathcal{L_{\textrm{edit}}}(\mathcal{F},{\bm{z}}^{(s)}_{k},{\bm{z}}^{(c)}_{k}) over the network parameters and learned codes.

Our first goal is to conduct the edit accurately – the edited radiance field should render views of the instance that reflect the user’s desired change. Our second goal is to conduct the edit efficiently. Editing a radiance field is time-consuming, as modifying weights requires dozens of forward and backward calls. Instead, the user should receive interactive feedback on their edits. To achieve these two goals, we consider the following strategies for selecting which parameters to update during editing.

Update the shape and color codes. One approach to this problem is to only update the latent codes of the instance, as illustrated in Figure 2(a). While optimizing such few parameters leads to a relatively efficient edit, as we will show, this method results in a low-quality edit.

Update the entire network. Another approach is to update all weights of the network, shown in Figure 2(b). As we will show, this method is slow and can lead to unwanted changes in unedited regions of the instance.

Hybrid updates. Our proposed solution, shown in Figure 2(c), achieves both accuracy and efficiency by updating specific layers of the network. To reduce computation, we finetune the later layers of the network only. These choices speed up the optimization by only computing gradients over the later layers instead of over the entire network. When editing colors, we update only Frad\mathcal{F}_{\text{rad}} and z(c){\bm{z}}^{(c)} in the network, which reduces optimization time by 3.7×3.7\times over optimizing the whole network (from 972972 to 260260 seconds). When editing shape, we update only Ffuse\mathcal{F}_{\text{fuse}} and Fdens\mathcal{F}_{\text{dens}}, which reduces optimization time by 3.2×3.2\times (from 1,0811{,}081 to 342342 seconds).

In Section 4.3, we further quantify the tradeoff between edit accuracy and efficiency. To further reduce computation, we take two additional steps during editing.

Subsampling user constraints. During training, we sample a small subset of user-specified rays. We find that this choice allows optimization to converge faster, as the problem size becomes smaller. For editing color, we randomly sample 6464 rays and for editing shape, we randomly sample a subset of 8,1928{,}192 rays. With this method, we obtain 24×24\times speedups for color edits and 2.9×2.9\times speedups for shape edits. Furthermore, we find that subsampling user constraints preserves edit quality; please refer to the appendix for additional discussion.

Feature caching. NeRF rendering can be slow, especially when the rendered views are high-resolution. To optimize view rendering during color edits, we cache the outputs of the network that are unchanged during the edit. Because we only optimize Frad\mathcal{F}_{\text{rad}} during color edits, the input to Frad\mathcal{F}_{\text{rad}} is unchanged during editing. Therefore, we cache the input features for each of the views displayed to the user to avoid unnecessary computation. This optimization reduces the rendering time for a 256×256256\times 256 image by 7.8×7.8\times (from 6.26.2 to under 0.80.8 seconds).

We also apply feature caching during optimization for shape and color edits. Similarly, we cache the outputs of the network that are unchanged during the optimization process to avoid unnecessary computation. Because the set of training rays is small during optimization, this caching is computationally feasible. We accelerate color edits by 3.2×3.2\times and shape edits by 1.9×1.9\times.

3 Color Editing Loss

In this section, we describe how to perform color edits with our conditional radiance field representation. To edit the color of a shape instance’s part, the user selects a desired color and scribbles a foreground mask over a rendered view indicating where the color should be applied. The user may optionally also scribble a background mask where the color should remain unchanged. These masks do not need to be detailed; instead, a few coarse scribbles for each mask suffice. The user provides these inputs through a user interface, which we discuss in the appendix. Given the desired target color and foreground/background masks, we seek to update the neural network F\mathcal{F} and the latent color vector z(c){\bm{z}}^{(c)} for the object instance to respect the user constraints.

Let cf{\bm{c}}_{f} be the desired color for a ray r{\bm{r}} at a pixel location within the foreground mask provided by the user scribble and let yf={(r,cf)}y_{f}=\left\{({\bm{r}},{\bm{c}}_{f})\right\} be the set of ray color pairs provided by the entire user scribble. Furthermore, for a ray r{\bm{r}} at a pixel location in the background mask, let cb{\bm{c}}_{b} be the original rendered color at the ray location. Let yb={(r,cb)}y_{b}=\left\{({\bm{r}},{\bm{c}}_{b})\right\} be the set of rays and colors provided by the background user scribble.

Given the user edit inputs (yf,yb)\left(y_{f},y_{b}\right), we define our reconstruction loss as the sum of squared-Euclidean distances between the output colors from the compositing operation C^\hat{C} to the target foreground and background colors:

Furthermore, we define a regularization term Lreg\mathcal{L}_{\textrm{reg}} to discourage large deviations from the original model by penalizing the squared difference between original and updated model weights.

We define our color editing loss as the sum of our reconstruction loss and our regularization loss

We optimize this loss over the latent color vector z(c){\bm{z}}^{(c)} and Frad\mathcal{F}_{\text{rad}} with λreg=10\lambda_{\textrm{reg}}=10.

4 Shape Editing Loss

For editing shapes, we describe two operations – shape part removal and shape part addition, which we outline next.

Shape part removal. To remove a shape part, the user scribbles over the desired removal region in a rendered view via the user interface. We take the scribbled regions of the view to be the foreground mask, and the non-scribbled regions of the view as the background mask. To construct the editing example, we whiten out the regions corresponding to the foreground mask.

Given the editing example, we optimize a density-based loss that encourages the inferred densities to be sparse. Let σr\sigma_{\bm{r}} be a vector of inferred density values for sampled points along a ray r{\bm{r}} at a pixel location and let yfy_{f} be the foreground set of rays for the entire user scribble.

We define the density loss Ldens\mathcal{L}_{\textrm{dens}} as the sum of entropies of the predicted density vectors σr\sigma_{{\bm{r}}} at foreground ray locations r{\bm{r}},

where we normalize all density vectors to be unit length. Penalizing the entropy along each ray encourages the inferred densities to be sparse, causing the model to predict zero density on the removed regions.

We define our shape removal loss as the sum of our reconstruction, density, and our regularization losses

We optimize this loss over Fdens\mathcal{F}_{\text{dens}} and Ffuse\mathcal{F}_{\text{fuse}} with λdens=0.01\lambda_{\textrm{dens}}=0.01 and λreg=10\lambda_{\textrm{reg}}=10.

The above method of obtaining the editing example assumes that the desired object part to remove does not occlude any other object part. We describe an additional slower method for obtaining the editing example which deals with occlusions in the appendix.

Shape part addition. To add a local part to a shape instance, we fit our network to a composite image comprising a region from a new object pasted into the original. To achieve this, the user first selects a original rendered view to edit. Our interface displays different instances under the same viewpoint and the user selects a new instance from which to copy. Then, the user copies a local region in the new instance by scribbling on the selected view. Finally, the user scribbles in the original view to select the desired paste location. For a ray in the paste location in the modified view, we render its color by using the shape code from the new instance and the color code from the original instance. We denote the modified regions of the composite view as the foreground region, and the unmodified regions as the background region.

We define our shape addition loss as the sum of our reconstruction and our regularization losses

and optimize over Fdens\mathcal{F}_{\text{dens}} and Ffuse\mathcal{F}_{\text{fuse}} with λreg=10\lambda_{\textrm{reg}}=10.

We note that this shape addition method can be slow due to the large number of training iterations. In the appendix, we describe a faster but less effective method which encourages inferred densities to match the copied densities.

Please refer to our video to see our editing demo in action.

Experiments

In this section, we show the qualitative and quantitative results of our approach, perform model ablations, and compare our method to several baselines.

Datasets. We demonstrate our method on three publicly available datasets of varying complexity: chairs from the PhotoShape dataset (large appearance variation), chairs from the Aubry chairs dataset (large shape variation), and cars from the GRAF CARLA dataset (single view per instance). For the PhotoShape dataset, we use 100100 instances with 4040 training views per instance. For the Aubry chairs dataset, we use 500500 instances with 3636 training views per instance. For the CARLA dataset, we use 1,0001{,}000 instances and have access to only a single training view per instance. For this dataset, to encourage color consistency across views, we regularize the view direction dependence of radiance, which we further study in the appendix. Furthermore, due to having access to only one view per instance, we forgo quantitative evaluation on the CARLA dataset and instead provide a qualitative evaluation.

Implementation details. Our shared shape network, instance-specific shape network, and fusion shape networks Fshare,Finst,Ffuse\mathcal{F}_{\text{share}},\mathcal{F}_{\text{inst}},\mathcal{F}_{\text{fuse}} are all 44 layers deep, 256256 channels wide MLPs with ReLU activations and outputs 256256 dimensional features. The shape and color codes are both 3232-dimensional and jointly optimized with the conditional radiance field model using the Adam optimizer and a learning rate of 10−410^{-4}. For each edit, we use Adam to optimize the parameters with a learning rate of 10−210^{-2}. Additional implementation details are included in the appendix.

Our method accurately models the shape and appearance differences across instances. To quantify this, we train our conditional radiance field on the PhotoShapes and Aubry chairs datasets and evaluate the rendering accuracy on held-out views over each instance. In Table 1, we measure the rendering quality with two metrics: PSNR and LPIPS . In the appendix, we provide additional evaluation using the SSIM metric in Table 4 and visualize reconstruction results in Figures 16-20. We find our model renders realistic views of each instance and, on the PhotoShapes dataset, matches the performance of training independent NeRF models for each instance.

We report an ablation study over the architectural choices of our method in Table 1. First, we train a standard NeRF over each dataset (Row 1). Then, we add a 6464-dimensional learned code for each instance to the standard NeRF and jointly train the code and the NeRF (Row 2). The learned codes are injected wherever positional or directional embeddings are injected in the original NeRF model. While this choice is able to model the shape and appearance differences across the instances, we find that adding separate shape and color codes for each instance (Row 3) and further using a shared shape branch (Row 4) improves performance. Finally, we report performance when training independent NeRF models on each instance separately (Row 5). In these experiments, we increase the width of the layers in the ablations to keep the number of parameters approximately equal across experiments. Notice how our conditional radiance network outperforms all ablations.

Moreover, we find that our method scales well to more training instances. When training with all 626626 instances of the PhotoShape dataset, our method achieves reconstruction PSNR 35.7935.79. We find that the shared shape branch helps our model scale to more instances. In contrast, a model trained without the shared shape branch achieves PSNR 33.9133.91.

2 Color Edits

Our method both propagates edits to the desired regions of the instance and generalizes to unseen views of the instance. We show several example color edits in Figure 3. To evaluate our choice of optimization parameters, we conduct an ablation study to quantify our edit quality.

For a source PhotoShapes training instance, we first find an unseen target instance in the PhotoShapes chair dataset with an identical shape but a different color. Our goal is to edit the source training instance to match the target instance across all viewpoints. We conduct three edits and show visual results on two: changing the color of a seat from brown to red (Edit 1), and darkening the seat and turning the back green (Edit 2). The details and results of the last edit can be found in the appendix. After each edit, we render 4040 views from the ground truth instance and the edited model, and quantify the difference. The averaged results over the three edits are summarized in Table 2.

We find that finetuning only the color code is unable to fit the desired edit. On the other hand, changing the entire network leads to large changes in the shape of the instance, as finetuning the earlier layers of the network can affect the downstream density output.

Next, we compare our method against two baseline methods: editing a single-instance NeRF and editing a GAN.

Single-instance NeRF baseline. We train a NeRF to model the source instance we would like to edit, and then apply our editing method to the single instance NeRF. The single instance NeRF shares the same architecture as our model.

GAN editing baselines. We also compare our method to the 2D GAN-based editing method based on Model Rewriting . We first train a StyleGAN2 model on the images of the PhotoShapes dataset . Then, we project unedited test views of the source instance into latent and noise vectors, using the StyleGAN2 projection method . Next, we invert the source and target view into its latent and noise vectors. With these image/latent pairs, we follow the method of Bau et al. and optimize the network to paste the regions of the target view onto the source view. After the optimization is complete, we feed the test set latent and noise vectors into the edited model to obtain edited views of our instance. In the appendix, we provide an additional comparison against naive finetuning of the whole generator.

These results are visualized in Figure 3 and in Table 2. A single-instance NeRF is unable to find an change in the model that generalizes to other views, due to the lack of category-specific appearance prior. Finetuning the model can lead to artifacts in other views of the model and can lead to color inconsistencies across views. Furthermore, 2D GAN-based editing methods fail to correctly modify the color of the object or maintain shape consistency across views, due to the lack of 3D representation.

3 Shape Edits

Our method is also able to learn to edit the shape of an instance and propagate the edit to unseen views. We show several shape editing examples in Figure 4. Similar to our analysis of color edits, we evaluate our choice of weights to optimize. For a source Aubry chair dataset training instance, we find an unseen target instance with a similar shape. We then conduct an edit to change the shape of the source instance to the target instance, and quantify the difference between the rendered and ground truth views. The averaged results across three edits are summarized in Table 3 and results of one edit are visualized in the top of Figure 4.

We find that the approaches of only optimizing the shape code and only optimizing Fdens\mathcal{F}_{\text{dens}} are unable to fit the desired edit, and instead leave the chair mostly unchanged. Optimizing the whole network leads to removal of the object part, but causes unwanted artifacts in the rest of the object. Instead, our method correctly removes the arms and fills the hole of the chairs, and generalizes this edit to unseen views of each instance.

4 Shape/Color Code Swapping

Our model succeeds in disentangling shape and color. When we change the color code input to the conditional radiance field while keeping the shape code unchanged, the resulting rendered views remain consistent in shape. Our model architecture enforces this consistency, as the density output that governs the shape of the instance is independent of the color code.

When changing the shape code input of the conditional radiance field while keeping the color code unchanged, the rendered views remain consistent in color. This is surprising because in our architecture, the radiance of a point is a function of both the shape code and the color code. Instead, the model has learned to disentangle color from shape when predicting radiance. These properties let us freely swap around shape and color codes, allowing for the transfer of shape and appearance across instances; we visualize this in Figure 5.

5 Real Image Editing

We demonstrate how to infer and edit extrapolated novel views for a single real image given a trained conditional radiance field. We assume that the single image has attributes similar to the conditional radiance field’s training data (e.g., object class, background). First, we estimate the image’s viewpoint by manually selecting a training set image with similar object pose. In practice, we find that a perfect pose estimation is not required. With the posed input image, we finetune the conditional radiance field by optimizing the standard NeRF photometric loss with respect to the image. When conducting this optimization, we first optimize the shape and color codes of the model, while keeping the MLP weights fixed, and then optimize all the parameters jointly. This optimization is more stable than the alternative of optimizing all parameters jointly from the start. Given the finetuned radiance field, we proceed with our editing methods to edit the shape and color of the instance. We demonstrate our results of editing a real photograph in Figure 6.

Discussion

We have introduced an approach for learning conditional radiance fields from a collection of 3D objects. Furthermore, we have shown how to perform intuitive editing operations using our learned disentangled representation. One limitation of our method is the interactivity of shape editing. Currently, it takes over a minute for a user to get feedback on their shape edit. The bulk of the editing operation computation is spent on rendering views, rather than editing itself. We are optimistic that NeRF rendering time improvements will help . Another limitation is our method fails to reconstruct novel object instances that are very different from other class instances. Despite these limitations, our approach opens up new avenues for exploring other advanced editing operations, such as relighting and changing an object’s physical properties for animation.

Acknowledgments. Part of this work while SL was an intern at Adobe Research. We would like to thank William T. Freeman for helpful discussions.

References

Appendix A Additional Experimental Details

Dataset rendering details. For the PhotoShape dataset , we use Blender to render 4040 views for each instance with a clean background. To obtain the clean background, we obtain the occupancy mask and set everywhere else to white. For the Aubry chairs dataset , we resize the images from 600×600600\times 600 resolution to 400×400400\times 400 resolution and take a 256×256256\times 256 center crop. For the CARLA dataset, we train our radiance field on exactly the same dataset as GRAF .

Conditional radiance field training details. As in Mildenhall et al. , we train two networks, a coarse network Fcoarse\mathcal{F}_{\textrm{coarse}} to estimate the density along the ray, and a fine network Ffine\mathcal{F}_{\textrm{fine}} that renders the rays at test time. We use stratified sampling to sample points for the coarse network and sample a hierarchical volume using the coarse network’s density outputs. The rendered outputs of these networks for an input ray r{\bm{r}} are given by C^coarse(r)\hat{C}_{\textrm{coarse}}({\bm{r}}) and C^fine(r)\hat{C}_{\textrm{fine}}({\bm{r}}), respectively. The networks are jointly trained with the shape and color codes to optimize a photometric loss. During each training iteration, we first sample an object instance k∈{1,…,K}k\in\{1,\dots,K\} and obtain the corresponding shape code zk(s){\bm{z}}^{(s)}_{k} and color code zk(c){\bm{z}}^{(c)}_{k}. Then, we sample a batch of training rays from the set of all rays RkR_{k} belonging to the instance kk, and optimize both networks using a photometric loss, which is the sum of squared-Euclidean distances between the predicted colors and ground truth colors,

When training all radiance field models, we optimize our parameters using Adam with a learning rate of 10−410^{-4}, β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, and ϵ=10−8{\epsilon}=10^{-8}. We train our models until convergence, which on average is around 11M iterations.

Conditional radiance field editing details. During editing, we keep the coarse network fixed and edit the fine network Ffine\mathcal{F}_{\textrm{fine}} only. We increase the learning rate to 10−210^{-2}, and we optimize network components and codes for 100100 iterations, keeping all other hyperparameters the same. To obtain training rays, we first randomly select an edited view of the instance, then randomly sample batches of rays from the subsampled set of foreground and background rays.

Conditional radiance field architecture details. Like the original NeRF paper , we use skip connections in our architecture: the shape code and embedded input points are fed into the fusion shape network as well as the very beginning of the network.

Furthermore, in our model architecture, we introduce a bottleneck that allows feature caching to be computationally feasible. The input to the color branch is an 88-dimensional vector, which we cache during color editing.

Last, when injecting a shape or color code to a layer, we run the code through a linear layer with ReLU nonlinearity and concatenate the output with the input to the layer.

GAN editing details. In our GAN experiments, we use the default StyleGAN2 configuration on the PhotoShapes dataset and train for 300,000300{,}000 iterations. For model rewriting , we optimize the 88-th layer for 2,0002{,}000 iterations with a learning rate of 10−210^{-2}.

On the CARLA dataset , we find that having only one training view per instance can cause color inconsistency across rendered views. To address this, we regularize the view dependence of the radiance c{\bm{c}} with an additional self-consistency loss that encourages the model to predict for a point x{\bm{x}} similar radiance across viewing directions d{\bm{d}}. This loss penalizes, for point x{\bm{x}} and viewing direction d{\bm{d}} in the radiance field, the squared difference between the sampled radiance c(x,d){\bm{c}}({\bm{x}},{\bm{d}}) and the average radiance of the point x{\bm{x}} over all viewing directions. Specifically, given radiance field inputs x,d{\bm{x}},{\bm{d}}, we minimize

In our experiments, we use K=64K=64. We visualize results with and without this regularization in Figure 7.

Appendix B Additional Evaluations

SSIM metric. We report additional evaluation on model ablation, color editing, and shape editing results using SSIM . Quantitative results can be found in this appendix’s Table 4 (model ablation), Table 5 (color edits), and Table 6 (shape edits).

Subsampling user constraints. During editing, we do not train on the whole foreground and background regions provided by the user, which can potentially decrease the quality of our edits. This is because training on fewer rays can cause the edit to propagate onto unwanted areas. For example, regions which the user specify as background, but are not in the set of sampled rays, can potentially be changed. However, we find that upon adding this optimization, the average PSNR over the three color edits decreases to 34.4934.49.

Additional GAN editing baselines. We compare our editing method against a naive generator fine-tuning method . The method is identical to the model rewriting method, except instead of conducting a low-rank update of the weights of a particular layer, we freely optimize all the weights of the generator. This optimization is done over 10,00010{,}000 steps with a learning rate of 10−310^{-3}. We report our results in Table 5.

Appendix C Additional Shape Editing Methods

Shape removal. For shape removal method described in the main paper, we assume that there is nothing behind an object part that a user scribbles over, allowing us to replace the object part with a white background. However, in practice, a user may wish to remove an object part that is in front of another one. To handle such occlusion, we propose a separate procedure: for each ray in the foreground mask, we zero out the first mode of density along the ray. We define the first mode of density to start at the first point with nonzero-density up to the first subsequent point with zero density. We find that this procedure is effective but can be slow and may leave artifacts of incomplete removal.

Shape addition. For shape addition, our method for reconstructing a composite image leads to effective but slow edits. We propose an additional density-based loss which is faster but less effective in executing the edit. The method for obtaining the editing example is the same, but we now optimize a loss that encourages the density values in the modified regions of the composite view to match with the density values of the regions copied from.

Specifically, let yf={(r,σf)}y_{f}=\{({\bm{r}},\sigma_{f})\} be the set of rays and densities in the foreground mask and yb={(r,σb)}y_{b}=\{({\bm{r}},\sigma_{b})\} be the set of rays and densities in the background mask. Here, σf\sigma_{f} are density values of the rays copied from the new object instance, and σb\sigma_{b} represent density values of the rays from the original instance. Furthermore, let σr\sigma_{\bm{r}} be the densities predicted by our model for ray r{\bm{r}}. Again, densities are normalized to sum to one.

which encourages the predicted densities to match the target densities for the edited regions and be unchanged for the unedited regions.

Appendix D User Interface

For our user interface, the user first picks an object instance they would like to edit. Our UI then displays several rendered views of that instance, and the user picks one view to edit. The user can then edit the selected view on an editing panel.

We provide four types of user edits: color edits, shape removal, shape addition, and color/shape transfer. Next, we describe the user interactions for each type of edit.

Color edits. The user chooses the target color from a color palette. Then, the user specifies a foreground mask over the view by clicking the edit color button, selecting a brush color, and scribbling over parts of the object. Last, the user specifies the background mask by clicking the BG brush and scribbling over where they would like to keep the part unchanged.

Shape removal. The user clicks the remove shape button and scribbles over parts of the image they would like to remove.

Shape addition. The user clicks the add shape button and several instances to copy shape from will pop up. The user specifies a target instance they would like to copy from, and a view of that instance is shown. Then, the user scribbles over the object part they would like to copy, and clicks on the location of the source instance where they would like to paste.

Shape/Color transfer. The user clicks either the transfer color button or the transfer shape button and several instances to transfer color/shape from will pop up. Then, the user clicks a desired target instance to transfer color/shape information.

Once the user editing is done, the user will click the execute button to execute the desired edit. Our algorithm will then finetune the latent variables and network weights and update the renderings of the edited object. Please see our video demo for more details.

Appendix E Additional Color Edits

Quantitative color editing evaluation. In this section, we provide the visualizations of all three color edits used for evaluation in the main paper. Visually, we again see that editing a single-instance NeRF leads to visual artifacts and visual inconsistencies across views. Similarly, GAN-based methods are unable to learn an edit that generalizes across views, likely due to their lack of a 3D representation. We visualize the results of the three edits in Figure 8 and quantify them in Table 7.

In Figure 8, the first two rows visualize Edits 1 and 2 discussed in the main paper, and are identical to the visualizations in the main paper’s Figure 3. The last row of Figure 8 visualizes Edit 3, which changes the seat of a chair from brown to green, then the chair back from beige to grey.

In Table 7, we quantify the quality of each of the four edits. The first two main columns correspond to the Edits 1 and 2 discussed in the main paper, while the last two main columns correspond to Edits 3 discussed in Section E.

Single-instance NeRF editing. We also provide an additional comparison of our method against editing a single-instance NeRF. Here, we change the color of a seat from brown to bright red. Again, we observe that the single-instance NeRF does not learn an edit that generalizes; the model frequently creates red artifacts in chair’s background. In contrast, our model can still learn an edit that successfully propagates to the seat but not to other regions of the scene. We visualize these results in Figure 9.

Additional color editing results. We visualize additional color edits on the Aubry chairs and the CARLA cars datasets in Figure 10.

Appendix F Additional Shape Edits

Quantitative color editing evaluation. In this section, we visualize all three shape edits used for evaluation in the main paper. We visualize the results of the three edits in Figure 11 and quantify them in Table 8. Visually, we see that consistent with the main paper, both finetuning the shape code and the shape branch are not enough to change the instance, but finetuning the whole network causes unwanted changes in the instance.

We compare our method against editing a single-instance NeRF . We find that similar to the case with color edits, single-instance NeRFs are unable to learn an edit that generalizes to unseen views, likely due to a lack of a category-level prior. We visualize these results on the PhotoShapes dataset in Figure 12.

Additional shape editing results. We visualize additional shape edits on the PhotoShapes and the CARLA cars datasets in Figure 13.

Appendix G Additional Color/Shape Swapping Edits

We visualize additional shape and color swapping results on the PhotoShapes dataset in Figure 14 (this appendix). Notice again how changing the color code keeps the shape of the instance unchanged, and how changing the shape code keeps the color of the instance unchanged.

Appendix H View Reconstruction

View consistency results. For each of our three datasets, we visualize synthesized views for a fixed instance and observe that the rendered views are all consistent in shape and color. We visualize these results in Figure 15. Notice how in the CARLA dataset , despite training on only one image per car instance, the model is able to infer the occluded regions of the car.

Additional reconstruction results. We visualize reconstructed views and depth maps across several instances of the PhotoShapes dataset using our conditional radiance field. For each instance, we render four unseen viewpoints from our model and visually compare them against the ground truth views. We find that our method is able to almost perfectly reconstruct each instance, as well as learn convincing depth estimates of each instance. We visualize reconstructions and depth maps for unseen views in Figures 16-20.

Appendix I Changelog

v2 Update Figure 8 and include additional details in the appendix.