Panoptic Neural Fields: A Semantic Object-Aware Neural Scene Representation
Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas Guibas, Andrea Tagliasacchi, Frank Dellaert, Thomas Funkhouser
Introduction
The ability to understand the content within an image is an essential task in computer vision, and over time we have witnessed a rapid increase in task complexity. Over a short period of time, we have progressed from the task of identifying the overall presence of objects within an image (i.e. classification and object detection ), to fine grained pixel-by-pixel classification (i.e. semantic segmentation ), and to the ability to differentiate between object instances of the same class (i.e. panoptic segmentation ).
However image level representations described above have limited applications. Instead we are interested in full 3D scene understanding which is important for autonomous driving , semantic mapping , and many other applications involving navigation or operation in the physical world . Given a sequence of RGB images, our goal is to infer: 1) a 3D reconstruction of the observed geometry, 2) a radiance field of the scene, 3) a decomposition of the scene into potentially dynamic things (e.g., cars) and background stuff (e.g., grass), 4) a category and instance label for every 3D point, as illustrated in Figure Panoptic Neural Fields: A Semantic Object-Aware Neural Scene Representation.
In recent years, neural 3D scene representations like NeRF have made significant advancements . NeRF represents a scene using a multi-layer perceptron (MLP) that maps positions and directions to densities and radiances which can then be used to synthesize an image from a novel view. However NeRF lacks semantic understanding and is also not object aware. In this work we explore neural scene representations for semantic 3D scene understanding tasks beyond the usual view synthesis task.
Some recent work augments NeRF to infer semantics , adding an extra head to predict semantic logits for any 3D position along with the usual density/color. Other recent work decomposes a scene into a set of NeRFs associated with foreground objects separated from the background . However, these systems have several limitations in the context of our goals: 1) they do not produce panoptic segmentations, 2) they learn from scratch for every scene; and 3) they share MLPs for multiple objects, which limits their ability to reproduce specific instances.
We address these issues in our proposed Panoptic Neural Fields (PNF), an object-aware neural scene representation that explicitly decomposes a scene into a set of objects (things) and amorphous stuff background. Each object instance is represented by a separate MLP to evaluate the radiance field within the local domain of a potentially moving and semantically labeled 3D bounding box. The semantic-radiance field of the stuff background is also represented by a MLP which includes an additional semantic head. Together the stuff and things MLPs jointly define a panoptic-radiance field that describes the density, color, category, and instance label of any 3D point over time.
Our object aware representation makes it possible to describe scenes with multiple moving objects and also paves the way to incorporate constraints that objects of the same category have similar shape and appearance. Previous object-aware frameworks used a shared MLP with instance-specific latent codes to incorporate this prior. In our model, each object instance is represented by a separate MLP that is initialized with a category-specific prior using meta-learning. The separation of learning of object category priors via meta-learning makes it possible to represent instance-specific details with smaller MLPs, which speeds inference in scenes with many objects.
Given a collection of images captured from a scene, we employ off-the-shelf algorithms to predict camera parameters and 2D semantic segmentations for all images, plus a set of 3D object detections with 3D oriented bounding boxes and category labels . We initialize the weights of the MLPs for our panoptic neural field model either with object category-specific meta-learned initialization or simple biased initialization of density activation layers. We then jointly optimize the bounding box and MLP parameters to minimize analysis-by-synthesis style losses that measure differences in color and semantic images synthesized with volumetric rendering (as in NeRF ). Thus, our approach provides an unified framework for optimizing 3D shape, appearance, semantics, and object poses all from a set of color images.
We evaluate our method on several scene understanding and synthesis tasks using experiments on the KITTI and KITTI-360 dataset, including 3D panoptic reconstruction, and scene editing. The output panoptic-radiance field can also be used to synthesize 2D image-level outputs like semantic segmentation, panoptic segmentation, depth images, and colored images of both observed and novel views. We demonstrate the utility of the proposed method for these scene understanding tasks, as well as for novel-view synthesis method with movable scene components.
Our contributions can be summarized as follows:
We propose, to the best of our knowledge, the first method that can derive a panoptic-radiance field of complex dynamic 3D scenes from images alone.
Our single unified model achieves state-of-the-art quality across multiple tasks and benchmarks on KITTI and KITTI-360 datasets.
We incorporate object shape and appearance priors via category-specific meta-learned initialization. This allows our object MLPs to be much smaller and faster than previous object-aware representations.
We jointly optimize all (stuff and things) neural fields and object poses, allowing our method to cope with noisy object poses and image segmentations.
Related Work
The most relevant related work is summarized in Table 1 and can be broadly divided into three categories: (1) Learning based single image 3D semantic and/or instance segmentation, (2) Multi-view 3D reconstruction and segmentation methods, and (3) Neural fields.
Single image reconstruction and segmentation. 3D-RCNN and Mesh-RCNN takes as input a single RGB image and predicts 3D mesh and pose of object instances in the image. Total3DUnderstanding combines layout estimation, 3D object detection, and object mesh generation. More recently showed 3D panoptic reconstruction and segmentation from a single RGB image.
Multi-view reconstruction and segmentation: Incorporating semantics into SLAM and SfM systems has a long history . More recently, PanopticFusion is an incremental, online mapping approach that fuses a sequence of RGB-D images into a consistent panoptic segmentation. ATLAS reconstructs and labels the 3D geometry from multiple posed RGB images. However, ATLAS produces only semantic segmentations (without instances). Both PanopticFusion and ATLAS works only for static scenes, requires 3D supervision, and relies upon convolutions on a discrete voxel grid, which limits its resolution. Kimera takes a stereo sequence and does online reconstruction, meshing, and semantic labeling of the mesh using ground-truth labels, as a proxy for any 2D segmentation method. Dynamic Scene Graphs expands on that by inferring object instances, even dynamic ones in case of people. Both methods, though representing impressive systems, were only demonstrated in simulation and rely on ground truth semantic labels.
Neural radiance fields (NeRFs). This work builds upon NeRF , which represents a scene using a multi-layer perceptron (MLP) that maps positions and directions to densities and radiances. From that representation, novel views can be synthesized using volumetric rendering and compared to input views in a self-supervised optimization procedure to infer MLP weights for an observed scene. However, NeRF only works for static scenes and trains for hours (from scratch) for every set of input views.
NeRFs with semantics. Recent work has considered using neural representations to infer semantics . In particular, SemanticNeRF adds an extra head to NeRF to predict semantic labels for any 3D position along with the usual density and color. Concurrent to our work also demonstrated neural panoptic fusion from multiple views. However both of these work are not object aware and cannot handle dynamic scenes.
NeRFs with dynamics disentangle a scene into a canonical volume and its time-varying deformation, represented by a second MLP. This approach has been applied for deforming faces , moving human bodies , and objects . In contrast, we consider dynamic scenes that contain many moving objects.
NeRFs with object decompositions decompose a scene into a set of NeRFs associated with foreground objects separated from the background. ObjectNeRF uses the object branch to render rays with masked areas for foreground objects conditioned on a latent code. Similarly, Neural Scene Graphs (NSGs) uses a separate conditional NeRF for each object category, and a multiplane neural representation for the background. However, these systems have several limitations in the context of our goals: 1) they do not produce panoptic segmentations, 2) they learn from scratch for every scene; and 3) they share NeRFs for multiple objects (which limits the ability to reproduce specific discrete instances).
Conditional NeRFs infer latent codes as pioneered in GRAF , piGAN , and PixelNeRF , as well as the recent CodeNeRF , which also optimizes over object poses. All these works incorporate category-specific priors by sharing MLP weights across object instances, combined with instance specific codes. We instead use instance-specific MLPs for representing each object, which allows each MLP to be smaller, resulting in faster inference speed on scenes with multiple objects. Object appearance and shape priors are incorporated via category-specific meta-learned initialization of the MLP weights.
Method
This section introduces the panoptic neural field representation and our computational pipeline (Figure 2). In Sec. 3.1 we describe the representation itself, which stores a panoptic-radiance field that can be used to query the color, density, semantic and instance labels at any 3D point at any time. In Sec. 3.2, we describe how this panoptic-radiance field can be rendered using NeRF-style volume rendering by over compositing along sampling of points along rays. In Sec. 3.3 we explain how the model is trained in analysis-by-synthesis style by comparing the rendered color and semantic segmentation with observed 2D color and predicted 2D semantic labels.
A key difference between our framework and previous object-aware frameworks is how we train and represent things. As illustrated in Fig. 3, our framework uses instance-specific fully weight encoded functions to represent each object, in comparison to the traditional approach of using a shared MLP with instance-specific latent codes. This design choice is driven by several factors. First, since the MLP only needs to represent a single object instance, we can have a smaller MLP compared to shared MLPs, resulting in faster inference speed on scenes with multiple objects. Second, this allows the object MLP to use its full capacity to describe and overfit to a specific novel object instance, which may not be possible with latent encoding. Third, it is simpler and does not require any change to the core NeRF model architecture. Object-level priors can also be incorporated to our instance-specific models using meta-learning based initialization (See Sec. 3.4).
Things: Foreground objects in our representation are represented by a neural function inside a dynamic bounding box. To instantiate the set of object tracks in a scene, we first run an RGB-only 3D object detector and tracker . This provides a bounding box track and semantic class for each recognized object instance . The track is parameterized by a sequence of transformation matrices, one at each of at a set of discrete timestamps. For each timestamp, we create a rotation matrix and a translation vector . There is also a box extent along each axis that is time-invariant. To determine the coordinate frame of an object at an arbitrary real-valued timestamp, we interpolate the discrete track steps.
Stuff: We represent the static background stuff with a single neural function. In addition to predicting density and color at every 3D point, the stuff function also learns a semantic label per point. We again use an MLP to represent the learned function. The architecture is similar to NeRF, but with an additional head for semantic logits. This head is direction-invariant to encode the inductive bias that 3D points have multi-view consistent semantic labels. Note that unlike the MLPs for objects, which are bounded, the stuff MLP must handle the unbounded nature of real world scenes. Therefore for large scenes, we follow and use a separate foreground and background stuff MLPs.
Panoptic-Radiance Field: The final panoptic-radiance field at a 3D point is computed from aggregating the contributions of the individual thing and stuff MLPs. For any given output channel (color, density, etc.), our function takes the sum of all contributions from any bounding box hits, defaulting to the stuff output if there is no intersection. For the color field , this is:
2 Rendering Panoptic-Radiance Fields
Given the complete panoptic-radiance field representation, a 2D image can be synthesized with volume rendering. This process is described in more detail in NeRF . Our image synthesis approach is the similar to NeRF, with the addition of support for extra output channels and dynamic boxes. To render a single ray we uniformly sample points (with jitter) along the ray and alpha-composite the result of the output channel we wish to render (RGB, depth, semantic, instance):
Above, is the final weight associated with each sample, determined by over compositing the opacity values of each sample along the ray. The function returns the representation value for the channel in question at the query point. For semantics, this is logits, while for instance it is a one-hot encoding of the object instance identifier .
3 Model Losses and training
We jointly optimize all network parameters and object tracks to reproduce the observed RGB images and predicted 2D semantic images:
At each gradient descent step, we randomly sample mini-batches of rays. Our RGB loss is the mean squared error between the synthesized and ground truth color, summed over sampled rays as in NeRF . Our semantic loss is applied at the same pixel locations, and compares the synthesized semantics with the input 2D semantic segmentation prediction . For this loss, we apply a per-pixel softmax-cross entropy function rather than mean squared error.
4 Incorporating priors via initialization
One of the core benefits of an object aware approach, is the ability to incorporate inductive bias that objects instances within same category, often have similar 3D shape and appearance. One possible way to incorporate such priors is to have shared MLP weights across all object instances, combined with some instance specific codes. Our framework instead uses separate MLPs for representing each object instance. As illustrated in Fig. 3, this allows each MLP to be smaller as it only needs to represent a single object instance, resulting in faster inference speed on scenes with multiple objects. Object shape and appearance priors are instead incorporated via initialization of the MLP weights of the neural functions. We present two approaches (see Fig. 4) of initializing our model, one based on category specific meta-learning and another based on simple bias initialization of the activation function of the MLPs.
Biased initialization: This simple initialization scheme improves convergence behavior and training performance without requiring a large dataset from which to learn a shape prior. In real-world outdoor scenes, most of the stuff volume is empty space. By contrast, most of the volume inside each things object bounding box is non-empty. We incorporate this prior directly by biasing the density prediction layer of the MLPs. For the stuff MLP, we initialize our bias to , whereas for all thing MLPs, we initialize the final bias layer to . Furthermore, for all MLPs, we use the softplus activation for the fully connected layer predicting the density outputs . We have found that this simple injection of prior knowledge via initializing the bias values is quite effective and robust compared to random initialization, since dense content in stuff can suppress gradients from distant objects.
Category specific learned initialization: If a large shape collection is available for certain object categories, we can further improve on the bias initialization scheme. In particular, we use meta-learning to capture category-specific shape and appearance priors. This process is illustrated in Fig. 5. First, we meta-learn category-specific initial weights by pre-training on ShapeNet . Then, we use these weights to initialize the thing MLPs in our model when training on a novel scene.
To meta-learn a category-specific shape prior, we use the FedAvg algorithm. This algorithm is known to be equivalent to REPTILE meta-learning, used before for NeRF initialization in . To do one meta-step of FedAvg, we independently optimize a set of MLPs, each on a separate ShapeNet shape. We then average the model weights across all NeRF models, and start another meta-step. Fig. 5 visualizes the evolution of the learned initialization across several outer loops of FedAvg. In our experiments, we pre-train using the 2D rendered car images of ShapeNet to obtain the learned initialization model. When reconstructing a full scene, we then initialize the MLPs for each car instance track to this initialization.
Evaluations
We performed a series of experiments to evaluate our model on multiple computer vision tasks, including view synthesis, reconstruction, 2D panoptic segmentation, 2D depth prediction, and scene editing. See the supplemental video for comprehensive visualizations of our results. All experiments used either the KITTI , Virtual KITTI or the recent KITTI-360 datasets. These datasets involves difficult forward facing cameras in complex outdoor dynamic scenes. KITTI-360 is the first benchmark that evaluates both the task of synthesising color and appearance images from novel views. Our model outperforms every other method in that leaderboard for both tasks as shown in Tab. 2. Since scenes in KITTI-360 chosen for the tasks are all static, we also evaluate our model on dynamic scenes from the KITTI dataset. Rendered views from these dynamic scenes are shown in Fig. 6 and Fig. 7. Additional results are available in Appendix B. Below we evaluate our model for each task in more detail.
Novel View Synthesis: How well a particular representation describes a scene is reflected in the quality of rendered views. As shown in Tab. 2, color images rendered by our model achieves the best PSNR and is competitive with latest view synthesis models . Since the scenes used in KITTI-360 novel view synthesis task are all static, we attribute the improved performance to our model benefiting from separate object-aware MLPs and incorporation of category-level priors. To study the novel view synthesis capabilities of our model on dynamics scenes, we also experiment on several dynamic scenes from KITTI dataset as shown in leftmost columns of Fig. 6 and Fig. 7. Notice that the rendered color images accurately captures the moving vehicles in the scene. A quantitative analysis of synthesized colored images for dynamic KITTI scenes is available in Tab. 3, wherein we follow the same experimental setup as described in Sec. 5.2 of Ost et al. . As expected, our method significantly outperforms representations like SRN and NeRF which rely on static world assumption. Note that our method also outperforms NSG even though our instance specific object MLPs are much smaller ( fewer FLOPs) compared to those used in NSG . The improvement over NSG also demonstrates the advantage of incorporating category-specific priors derived from meta-learning with images from ShapeNet.
Panoptic Segmentation: Semantic and instance segmentation of an arbitrary view can be obtained from our model by simple rendering (see Eq. 2) along the desired view. The two right columns of Fig. 6 and Fig. 7 demonstrates the rendered semantic and instance segmentation images from our model. As shown in Tab. 2, our model achieves state-of-the-art 74.28 mIoU for the task of novel view semantic segmentation on KITTI-360. One approach of generating segmentation images for any arbitrary view is to first synthesize color image from a desired view using view synthesis methods like ; followed by 2D image segmentation . However as demonstrated in Tab. 2, our unified model significantly outperforms all such two-stage baselines. Moreover the rendered segmentation images from our model are temporally consistent and works for dynamic scenes. We also perform a ablation study of the rendered segmentation images on dynamic scenes from KITTI. Quantitative and qualitative results are shown in Tab. 4 and Fig. 8 respectively. Our model significantly out performs (+9.2 mIoU) non object-aware models like SemanticNeRF , since they cannot model dynamic objects. Our model also improves upon single image state-of-the-art segmentation models by fusing information from multiple views.
2D Depth Estimation: We also demonstrate rendered depth images from our model, obtained by over compositing point depths along rays with opacity values as described in Eq. 2. Rendered depth images are shown in the second columns of Figs. 6 and 7, along with other outputs of the model. Please note the sharp reconstruction of shapes of moving cars on the road. Of course, those cars would be missing or blurred in standard NeRF models. This demonstrates the benefits of our object-aware approach that handles dynamic scenes and instance-specific object MLPs that accurately captures each object instance.
Object Decomposition: Fig. 9 shows visualizations of how images from a dynamic KITTI scene are decomposed into MLPs representing different object instances. These images are rendered without the (stuff) background. Compared to NSG , our method does a better job of disentangling objects from the (stuff) background. In particular, note that the occluding traffic sign-posts in front of the car are entangled with the rendered cars in NSG, but not in our results. Better object decomposition from our model is also important for the scene editing task discussed below.
Scene Editing: Since our model separates objects from the background and builds a full 3D radiance field for each object, it is possible to edit images using the model by removing objects, adding new objects, and transforming object bounding boxes and poses. Fig. 10 shows few scene editing examples on Virtual KITTI dataset. The top row shows original images, and the bottom row shows edits. In bottom-left, we demonstrate cloning of cars by replicating the weights of all object MLPs to a same car. In bottom right of the figure we independently rotate each vehicle object.
Limitations
Like most other NeRF-style methods, our model is compute-intensive and hence currently only suited for offline applications. However, we expect advances in neural rendering will alleviate some of these speed issues in near future. It also does not incorporate more complex light transport effects, such as shadowing, under object motion. Our framework optimizes and corrects bounding box poses from noisy 3D object detection and tracking, but has not been designed to handle other errors such as missing and duplicate detections and incorrect class predictions. Finally, our framework does not handle deformable objects and is restricted to scenes with rigid moving objects.
Conclusion
This paper presents Panoptic Neural Fields (PNF), an object-aware neural scene representation that decomposes a scene into a set of MLPs associated with object instances (things) and the background (stuff). Our model learns a 4D panoptic radiance representation of dynamic scenes from images alone. This representation can be queried to obtain the color, density, instance, and category label of any 3D point over time. Several tasks like scene editing, view synthesis, panoptic segmentation are derived by simply rendering the representation from the desired views. Results of experiments on several KITTI scenes demonstrate state-of-the-art performance for novel view synthesis and panoptic segmentation for challenging outdoor scenes with multiple dynamic objects.
Appendix A Additional Model Details
In this section we provide additional details for training and inference of the proposed panoptic neural field model described in Sec. 3.
For stuff MLP, we use hidden layers of width . For thing MLP, we use MLP with hidden layers and width . We use positional encoding of frequencies to encode position coordinates for stuff MLP and frequencies to encode position coordinates in object coordinate space for each thing MLP.
The density and semantic output does not depend on view directions, whereas the color output is additionally conditioned on view directions similar to . This encodes the assumption that structure and semantics of the static background and individual objects only varies with position coordinates in their respective coordinate spaces. We use frequencies to encode the view directions before feeding them to the stuff and thing MLPs. For view directions, we find it beneficial to gradually activate the encoding frequencies over the course of the optimization, similar to .
To conserve memory we do not perform hierarchical sampling . Instead we sample additional points with stratified sampling strategy along rays while training and inference. So unlike which uses a pair of MLPs corresponding to coarse and fine sampling, only one set of MLPs are required in our model. We used 1024 samples per ray in our experiments on KITTI dataset.
A.2 Model Initialization Details
As described in Sec. 3.4, we incorporate priors via initialization. For cars and vans, we use category specific initialization of thing MLP weights. This category specific initialization is meta-learned from 2D rendered images of 3D cars models from ShapeNet dataset. For all other object categories, we use the biased initialization described in Sec. 3.4. We use a simplified federated averaging algorithm as described in Algorithm 1 to realize our category specific learned initialization. Also see Fig. 5 which visualizes the evolution of the learned initialization across several outer loops of FedAvg. This algorithm is known to be equivalent to REPTILE meta-learning. The main advantage of the simple FedAvg is that it allows decentralized federated training. In Fig. 5, we assumed simple SGD as the optimization algorithm, but can be easily adapted for other optimizers.
Appendix B Additional Results
In this section we provide additional results on KITTI and KITTI-360 for the tasks of novel view synthesis of color, semantics and depth images, along with scene editing. We also evaluate the benefits of our proposed category specific meta-learned initialization of thing MLP weights on the ShapeNet dataset.
Additional qualitative results of our model on KITTI and KITTI-360 are shown on Fig. 12 and Fig. 11 respectively. Just like our experiments in Sec. 4, we only used the forward facing cameras images for these results. Note that even though KITTI-360 dataset provides side facing fisheye camera images, only the forward facing stereo images are available for the novel view test sequences used for the experiments. Each row of Fig. 12 and Fig. 11 shows a novel view of a scene generated by rendering the panoptic neural field representation of that scene. The left columns shows the rendered semantic segmentation overlaid on top of the rendered color image so that it is easier to judge the segmentation quality. The colormap used for visualizing the segmentation and depth images are shown at the bottom. The corresponding rendered depth images from those views are shown in the right columns. Note that even the difficult thin structures like lamp poles and sign posts are both accurately reconstructed and segmented by our model. Also note the accurate reconstruction and segmentation of multiple moving cars in Fig. 12. Additionally, rendered color images from some novel viewpoints are also available the left column of Fig. 13.
B.2 Scene Editing
Since our proposed panoptic neural field scene representation is object aware, it allows seamless manipulation and editing of different objects present in the scene. In addition to the scene editing results on Virtual KITTI dataset discussed in Sec. 4 and Fig. 10, we show scene editing results on KITTI and KITTI-360 datasets in Fig. 13. Each row in Fig. 13 shows rendered color images along some novel view of both the original (left) and edited (right) scene representations. More specifically we show adding and removing new objects into the scene, changing the 3D pose of objects, and object cloning where the thing MLP parameters are replicated for all objects in the scene.
B.3 Benefits of our learned initialization
In real world scenes, objects are often captured from a sparse set of views. For example in self driving car scenes like KITTI, most objects (e.g. cars) often get a limited set of views from one side only. Thus incorporating prior knowledge becomes important for completeness and accurate reconstruction.
We learn the category specific priors from a large collection of objects on ShapeNet, as part of a separate meta-learning process and distill that knowledge as initialization when training on a novel scene. Thus our inference time scene representation network is more efficient and only focus on the individual set of object instances present in the scene. As demonstrated in Tab. 3, our model does a better job in reconstructing images of a dynamic scene compared to other object aware approaches like NSG , even though we use a much smaller MLP (10x fewer FLOPs) per object.
The learned initialization also provides other benefits like faster convergence and better completeness when reconstructed from sparse partial observations. We demonstrate these two benefits on ShapeNet dataset in Fig. 14 and Fig. 15. Specifically, we used rendered images of cars from ShapeNet provided by .
Fig. 14 qualitatively compares the rendered color images when using learned initialization over the standard xavier (glorot) initialization after two full epochs of training. For this experiment each model has at-least 50 input views. As seen in Fig. 14, the proposed initialization offers clear benefits in terms of faster convergence even when there is a dense set of input views.
The advantage of the learned initialization over standard initialization is more pronounced when we have few sparse input views of an object. This is demonstrated in Fig. 15. Using the proposed learned initialization, our model can reconstruct novel object instances even with just a single image as input. As shown in Fig. 15, even when only a partial view of the objects are used as input, the category specific object priors distilled via the initialization results in a more complete reconstruction.
B.4 Video Results
We also encourage the readers to also look at supplemental videos demonstrating results of our framework and a overview of the method. Most of our results on dynamic scenes are better visualized in the video.
Appendix C Potential Negative Societal Impact
Our contribution is an intermediate representation for comprehensive 3D scene understanding. We believe this can enable applications with a beneficial impact on society. However, it could also enable applications with potential negative impact. While it is impossible to anticipate all possible such applications, we discuss a few below.
Because our method supports comprehensive tracking of objects and people, it could be extended for use in crowd monitoring, traffic density reports and beneficial applications stemming from that. However, it could also be incorporated into surveillance systems. We will include stipulations in the license agreement for the code limiting its applications to academic research.
In addition, because our methods support view synthesis of 3D scenes, it is conceivable that it could be used to create imagery of fictional events, with the potential to disseminate fake news and/or propaganda. Because our method supports scene editing, actual events could be altered and used in similar ways. Of course, we will clearly mark all images generated by our system as “synthetic.” Additionally, we will include a requirement to do the same in the download instructions for our code.
Mitigation of the above issues is hard: many computer vision contributions are intermediate representations like ours. Segmentation, feature tracking, and object recognition can be put together into diverse functioning applications. As a profession, we should strive for the ethical application of these new technologies.