Learning Object-Centric Representations of Multi-Object Scenes from Multiple Views
Li Nanbo, Cian Eastwood, Robert B. Fisher
Introduction
Traditional VAEs use “single-object” or “flat” vector representations that fail to capture the compositional structure of natural scenes, i.e. the existence of interchangeable objects with common properties. As a result, “multi-object” or object-centric representations have emerged as a promising approach to scene understanding, improving sample-efficiency and generalization for many downstream applications like relational reasoning and control . However, recent progress in unsupervised object-centric scene representation has been limited to “single-view” methods which form their representations of 3D scenes based only on a single 2D observation (view). As a result, these methods form inaccurate representations that fall victim to single-view spatial ambiguities (e.g. occluded or partially occluded objects) and fail to capture 3D spatial structures.
To address this, we present MulMON (Multi-View and Multi-Object Network)—an unsupervised method for learning object-centric scene representations from multiple views. Using a spatial mixture model and iterative amortized inference , MulMON sidesteps the main technical difficulty of the multi-object-multi-view scenario—maintaining object correspondence across views—by iteratively updating the latent object representations for a scene over multiple views, each time using the previous iterations posterior as the new prior. To ensure that these iterative updates do indeed aggregate spatial information, rather than simply overwrite, MulMON is asked to predict the appearance of the scene from novel viewpoints during training. Given images of a static scene from several viewpoints, MulMON forms an object-centric representation, then uses this representation to predict the appearance and object segmentations of that scene from unobserved viewpoints. Through experiments we demonstrate that:
MulMON better-resolves spatial ambiguities than single-view methods like IODINE , while providing all the benefits of object-based representations that “single-object” methods like GQN lack, e.g. object segmentations and manipulations (see Section 5).
MulMON accurately captures 3D scene information (rotation along the vertical axis) by integrating spatial information from multiple views (see Section 5.3).
MulMON achieves both inter- and intra-object disentanglement—enabling both single-object and single-object-property scene manipulations (see Section 5.3).
MulMON represents the first feasible solution to the multi-object-multi-view problem, permitting new functionality like viewpoint-queried object-segmentation (see Section 5.2).
Background
2 Multi-object representations from multiple views
As 2D views of 3D scenes are inherently under-specified, they often contain spatial ambiguities (e.g. occlusions or partial occlusions). As a result, “single-view” methods often learn inaccurate representations of the underlying 3D scene. To solve this, it is desirable to aggregate information from multiple views into a single accurate representation of the (static) 3D scene. Recently, this was achieved with some success by GQN for single-object (i.e. single-slot) representations. However, doing so for multi-object representations is much more challenging due to the difficultly of the object matching problem, i.e. maintaining object-representation correspondences across views. As a result, this multi-object-multi-view (MOMV) problem remains unsolved.
Method
Our goal is to learn structured, object-centric scene representations that accurately capture the spatial structure of 3D multi-object scenes, and to do this by leveraging multiple 2D views. Key to achieving this is 1) an outer loop that iterates over views, aggregating information while avoiding the object matching problem, and 2) a training procedure that ensures that these outer loops are indeed used to form a complete 3D understanding of the scene, rather than just overwriting each other. We detail 1) in Section 3.1, and 2) in Section 3.4. Additionally, we describe the viewpoint-conditioned generative model and iterative inference procedure in Sections 3.2 and 3.3 respectively.
For a static scene, we consider that the latent scene representation is updated sequentially in steps as the observations are obtained one-by-one from to , where denotes the updating step. This suggests that is obtained by updating using a new observation , taken at viewpoint (see the green box in Figure 2(a)). Therefore, by making an assumption that for any integer , we can compute the target multi-view posterior in a recursive form as:
where is the latent representation of a scene object (indexed by ) before observing at viewpoint , the representation afterwards. and the initial guess which we assume to be a standard Gaussian distribution . The formulation in equation 1 turns the multi-view problem into a recursive single-view problem and, in theory, enables online learning of scenes from an infinitely large number of observations without causing memory overflow.
Within-view iterations (inner loop).
As shown in Figure 2, MulMON consists of a scene-representation inference model and a viewpoint-conditioned generative model. In each iteration, the inference model starts with a prior assumption about the objects in the latent space, i.e. , and approximates the target posterior after observing at viewpoint . The approximation, as mentioned in Section 1, is handled by iterative amortized inference and the approximate posterior is passed to the next iteration as the prior assumption. Therefore, a single iteration is a single-view process that takes a latent prior about and an image observation (taken at a viewpoint ) as inputs. We call the single-view iterative process the inner loop, and the cross-view Bayesian updating process (see equation 1) the outer loop.
2 Generative Model
We model an image with a spatial Gaussian mixture model, similar to MONet and IODINE, and additionally we take as input (condition on) the viewpoint . We can then write the generative likelihood as:
where are the RGB values in image at at a pixel location that pertain to object , is the Gaussian density function parametrized by a neural network , and is the mixing coefficient for object and pixel , i.e. the probability that pixel is assigned to the -th object. More formally, , where is a categorical random variable and represents the event that pixel is assigned to the -th object. This is an important property for object segmentation, as it implies that every pixel in must be explained by one and only one object. Together, the mixing coefficients for object (one per pixel) form a soft object segmentation mask . We assume all pixel values are independent given the corresponding latent object representation and viewpoint , and simplify computations by using a fixed variance for all pixels. In practice, we split the parameters into two pieces, and , in order to handle the viewpoint-queried neural transformation and observation-generation separately in two consecutive stages. That is, we first transform the latent object representations w.r.t. a viewpoint using the function , then we pass the output through a decoder in order to render a viewpoint-queried observation . We illustrate this process in Figure 2(b) and Algorithm 1.
3 Inference
Equation 1 simplifies the inference problem by breaking the computation of the multi-object-multi-view posterior , into a recursive computation of multi-object-single-view posteriors, . However, exact inference of is still intractable. Similar to IODINE, we apply iterative amortized inference to approximate the intractable target posterior. However, unlike IODINE which always initializes the prior from a standard Gaussian, the inference model of MulMON takes an approximate posterior from last iteration as the prior. Hence, we approximate the intractable posterior with , where parametrizes a set of object-specific Gaussian distributions in the latent space. We denote the number of iterations for the inner loop with , and each iteration is indexed by . The parameter update in the iterative inference is thus:
where the refinement function , with trainable parameter , is modeled by a recurrent neural network. The and operators denote parallel operations over independent object slots. The same auxiliary inputs such as mask gradients and posterior gradient , where is the objective function of MulMON (will be discussed in Section 3.4), as that of IODINE are also adopted to refine the posterior. These auxiliary inputs are computed by a function , namely the “auxiliary function”, which takes in the refinement function’s inputs along with the posterior parameter .
The variational approximate posterior of MulMON is thus:
where the initial guess is a standard Gaussian and this is the same as in equation 1. We refer to Algorithm 1 for more details about MulMON’s inference process and its behaviors at test time.
4 Training
Similar to IODINE, MulMON learns the decoder parameters and the refinement network parameters by minimizing , which is equivalent to maximizing the evidence lower bound (i.e. the ELBO, denoted as ) in a generative configuration. However, instead of directly maximizing the ELBO like IODINE, we simulate novel viewpoint-queried generation in the training process (similar to GQN ). By asking MulMON to predict the appearance of a scene from unobserved viewpoints during training, we ensure that the iterative updates are indeed used to aggregate spatial information across views, as a complete 3D scene understanding is required to perform well. More formally, we randomly partition the set of scene observations into two subsets and , with observations in and the remaining observations in . We perform scene learning on and novel viewpoint-queried generation on . We thus derive the MulMON ELBO (for one scene sample) as:
where is the information gain (aka. Bayesian surprise), the operation measures the size of a discrete set, and is an abbreviation of the variational posterior , from which we sample by applying ancestral sampling. In practice, we use an efficient approximation of the information gain, i.e. an approximate or . Note that using a fixed number of observations could harm the model’s robustness at test time, hence why we randomly partition the observations into size-varying sets and during training, i.e. we train the model with varying number of observations. See Appendix A for full details of the training algorithm of MulMON.
Related Work
Many recent breakthroughs in unsupervised representation learning of have come in the form of “disentanglement” models that seek feature-attribute-level understanding by encouraging e.g. independence among latent dimensions. However, most of these models focus on a single view of a single object that has been placed in front of some background (e.g. dSprites, CelebA, 3D Chairs). As a result, they fail to i) generalize to more realistic, multi-object scenes, and ii) accurately capture 3D scene information (e.g. resolve single-view spatial ambiguities and estimate e.g. rotation along the vertical axis).
Multi-object-single-view (MOSV).
To avoid the additional computational complexities of factorising or segmenting the objects in a scene into explicit multi-object representations, many works have used pre-segmented images . However, this comes at the cost of decreased representational power (good object representation requires good object segmentation ) and a reliance on annotated data. In addition, these works struggle in a multi-view scenario where pre-segmented images require consistent multi-frame object registration and tracking, since the segmentation and representation models work independently. More recently, several works have succeeded in approximating the factorized posterior within the VAE framework, achieving impressive unsupervised object-level scene factorization. However, being single-view models, they fall victim to single-view spatial ambiguities. As a result, they fail to accurately capture the scene’s 3D spatial structure, causing problems for object-level segmentation. To overcome this and learn object-based representations that accurately capture 3D spatial structures, MulMON essentially extends these models to the multi-view scenario.
Single-object-multi-view (SOMV).
Recently, unsupervised models like GQN and EGQN have been quite successful in aggregating multiple observations of a scene into a single-slot representation that accurately captures the spatial layout of the 3D scene, as shown by their ability to predict the appearance of a scene from unobserved viewpoints. However, being single-slot or “single-object” models, they fail to achieve object-level scene understanding in multi-object scenes, and as a result, miss out on the aforementioned benefits of object-centric scene representations. To overcome this, MulMON essentially extends these models to the case of multi-object representations. In addition, several works have sought explicit 3D representations either in the latent space or output space . However, due to the complexity of 1) working with explicit 3D object representations and 2) maintaining object correspondences across views, these works have been limited to single-object scenes (often quite simple, with “floating” objects placed in front of a plain background).
Multi-object scenes in videos.
While some works in multi-object discovery and tracking in videos appear to be MOMV models , they in fact work with one view per scene (abiding strictly by our definition of a scene in Section 2) and are only capable of dealing with binarizable MNIST-like images.
Experiments
Our experiments are designed to demonstrate that MulMON is a feasible solution to the MOMV problem, and to demonstrate that MulMON learns better representations than the MOSV and SOMV models by resolving spatial ambiguity. To do so, we compare the performance of MulMON against two baseline models, IODINE (MOSV) and GQN (SOMV), in terms of segmentation, viewpoint-queried prediction (appearance and segmentation) and disentanglement (inter- and intra-object). To best facilitate these comparisons, we created two new datasets called CLEVR6-MultiView (abbr. CLE-MV) and CLEVR6-Augmented (abbr. CLE-Aug) which contain ground-truth segmentation masks and shape descriptions (e.g. colors, materials, etc.). The CLE-MV dataset is a multi-view, observation-enabled variant (10 views per scene) of the CLEVR6 dataset. The CLE-Aug adds more complex shapes (e.g. horses, ducks, and teapots etc.) to the CLE-MV environment. In addition, we compare the models on the GQN-Jaco dataset and use the GQN-Shepard-Metzler7 dataset (abbr. Shep7) for a specific ablation study. We train all models using an Adam optimizer with an initial learning rate for gradient steps. In addition, all experiments were run across five different random seeds to simulate scenarios of different observation orders and view selections. For more details about the four datasets and model implementations see Appendix B and C respectively in the supplementary materials.
The ability of MulMON to perform scene object decomposition in the scene learning phase is crucial for learning object-centric scene representations. We evaluate its segmentation ability by computing mean-intersection-over-union (mIoU) scores between the output and the GT masks. However, since the segments produced by IODINE and MulMON are unordered, GT masks and object segmentation masks need to first be one-to-one registered for each scene. We solve this matching problem by first computing every possible object pairs of GT object masks and outputs, then, for each GT object mask, we find the output object mask that gives the highest IoU score. Table 1(a) shows that MulMON outperforms IODINE in object segmentation. The qualitative comparison in Figure 3 shows that IODINE captures each object well independently but fails to understand the spatial structure along depth directions (3D) – as described by the Categorical distribution (see Section 3.2). Note that IODINE’s poor segmentation performance is mostly due to its poor handling of the background, i.e. its tendency to split up the background. Although the background is often considered a less-important “object”, correct handling of the background demonstrates better spatial-reasoning ability. Together, all of these results suggest that MulMON learns better single-object representations and spatial structures by overcoming spatial ambiguities.
2 Novel-viewpoint Prediction
MulMON can predict both observations and segmentation for novel viewpoints. This is the major advantage of our model (a MOMV model) over the MOSV and SOMV models in scene understanding. For our evaluation of online scene learning, each model is provided with 5 observations of each scene and then asked to predict both the observation and segmentation for randomly-selected novel viewpoints. We compute the root-mean-square error (RMSE) and mIoU as quality measures of the predicted observation and segmentation respectively. Table 1(c) shows that MulMON outperforms GQN on novel-view observation prediction. Table 1(b) shows that MulMON is the only model that can predict the object segmentation for novel viewpoints – and it does so with a similar quality to the original object segmentation (compare with Table 1(a)). However, as shown in Figure 4, GQN tends to capture more pixel details than MulMON, albeit at the risk of predicting wrong spatial configurations.
3 Disentanglement Analysis
To evaluate how well MulMON performs disentanglement at both the inter-object level and the intra-object level, we run disentanglement analyses on the representations learned by MulMON. For our qualitative analysis, we pick one of objects in a scene, and traverse one dimension of the learned object-representation at a time. Figure 5 (left) shows i) MulMON’s intra-object disentanglement, encoding interpretable features in different latent dimensions; and ii) MulMON’s inter-object disentanglement, allowing single-object manipulation without affecting other objects in the scene. Figure 5 (right) shows that MulMON captures 3D information (vertical-axis rotation) and broadcasts consistent manipulations of this 3D information to different views. For our quantitative analysis, we employ the method of Eastwood and Williams to compare the representations learned by each model on the CLE-MV and CLE-Aug datasets. As shown in Table 1(d), MulMON learns object representations that are more disentangled, complete (compact) and informative (about ground-truth object properties). See Appendix D for further details.
4 Ablation Study
We consider the number of observations to be the most important hyperparameter of MulMON as the key insight of MulMON is to reduce multi-object spatial uncertainty by aggregating information across multiple observations. To visualize the effect of on MulMON’s performance, we plot MulMON’s uncertainty about the scene as a function of . More specifically, for a given scene and ordering of the observations, we: 1) draw samples from the approximate variational posterior , 2) obtain the corresponding viewpoint-queried observation predictions using the 10 latent samples (see Section 3.2 and Figure 2(b)), 3) compute the pixel-wise empirical variance over these observation predictions. Averaging over all scenes in the dataset and sampling 5 random view orderings (5 different random seeds), we can then create Figure 6 which shows that MulMON effectively reduces the spatial uncertainty/ambiguity by leveraging multiple views. In particular, MulMON’s uncertainty is rapidly reduced after only a small number of observations . We also study the effects of two other important hyperparameters, namely the globally-fixed number of object slots and the coefficient of information gain (in the MulMON ELBO). For details on these further ablation studies, we refer the reader to Appendix D.
Conclusion
We have presented MulMON—a method for learning accurate, object-centric representations of multi-object scenes by leveraging multiple views. We have shown that MulMON’s ability to aggregate information across multiple views does indeed allow it to better-resolve spatial ambiguity (or uncertainty) and better-capture 3D spatial structures, and as a result, outperform state-of-the-art models for unsupervised object segmentation. We have also shown that, by virtue of addressing the more complicated multi-object-multi-view scenario, MulMON achieves new functionality—the prediction of both appearance and object segmentations for novel viewpoints. As all scenes in this paper are static, future work may look to extend MulMON to dynamic multi-object scenes.
Acknowledgements
This research is partly supported by the Trimbot2020 project, which is funded by the European Union Horizon 2020 programme. The authors would like to thank Prof. C. K. I. Williams for his valuable advice that helps to improve this work and acknowledge the GPU computing support from Dr. Zhibin Li’s Advanced Intelligence Robotics Lab at the University of Edinburgh.
Broader Impact
In this paper, we presented a new method to learn object-centric representations of multi-object scenes. Object-centric scene representations can support many downstream tasks, such as autonomous scene exploration, object segmentation (tracking) and scene synthesizing.
Autonomous scene exploration has real-world applications in exploring hazardous environments, mines, potential bomb threats, nuclear waste zones. This could have societal impacts through increased worker safety or potential military (mis)uses.
Object detection and tracking has real-world applications in tracking people in CCTV footage, detecting buildings from aerial footage, and spotting potential hazards for autonomous vehicles. Potential societal impacts include safer autonomous vehicles and unwanted/increased surveillance.
Finally, scene synthesizing has applications in automated scene modelling for computer games. This further transitions society away from labor-intensive tasks to higher-level cognitive tasks. This could have both positive (more time for cognitive tasks) and negative (less employment) impacts on society.
References
A. Training Algorithm of MulMON
We refer to Algorithm 2 and 3 for the training algorithm of MulMON.
B. Data Configurations
We show samples of the used datasets in Figure 14.
CLEVR-MultiView & CLEVR-Augmented We adapt the Blender environment of the original CLEVR datasets to render both datasets. We make a scene by randomly sampling rigid shapes as well as their properties like poses, materials, colors etc.. For the CLEVR-MultiView (CLE-MV) dataset, we sample shapes from three categories: cubes, spheres, and cylinders, which are the same as the original CLEVR dataset. For the CLEVR-Augmented (CLE-Aug), we add more shape categories into the pool: mugs, teapots, ducks, owls, and horses. We render 10 image observations for each scene and save the 10 camera poses as 10 viewpoint vectors. We use resolution for the CLE-MV images and for the CLE-Aug images. All viewpoints are at the same horizontal level but different azimuth with their focuses locked at the scene center. We thus parametrize a viewpoint -D viewpoint vector as , where is the azimuth angle and is the distance to the scene center. In addition, we save the object properties (e.g. shape categories, materials, and colors etc.) and generate objects’ segmentation masks for quantitative evaluations. CLEVR-MultiView (CLE-MV) contains 1500 training scenes, 200 testing images. CLEVR-Augmented (CLE-Aug) contains 2434 training scenes and 500 testing scenes.
GQN-Jaco We use a mini subset of the original GQN-Jaco dataset in our paper. The original GQN-Jaco contains million scenes, each of them contains 20 image observations (resolution: ) and 20 corresponding viewpoint vectors (D). To reduce the storage memory and accelerate the training, we randomly sample scenes for training and scenes for testing. Also, for each scene, we use only 11 observations (viewpoints) that are randomly sampled from the 20 observations of the original dataset.
GQN-Shepard-Metzler7 Same as the GQN-Jaco dataset, we make a mini GQN-Shepard-Metzler7 dataset (Shep7) by randomly selecting 3000 scenes for training and 200 for testing. Each scene contains 15 images observations (resolution: ) with 15 corresponding viewpoint vectors (D). We use Shep7 to study the effect of on our model.
C. Implementation Details
Training configurations See Table 2 for our training configurations.
Model architecture We show our model configurations in Table 3, C. Implementation Details, and C. Implementation Details.
For a single view of a scene, our decoder outputs K RGB values (i.e. as in equation 2 of the main paper) along with K mask logits (denoted as ). and are the same as the image sizes, i.e. height and width. In this section, we detail the computation of rendering K individual scene components’ images, segmentation masks, and reconstructed scene images. We compute the individual scene objects’ images as:
As shown in Figure 8, this overcomes mutual occlusions of the objects since the functions do not impose any dependence on K objects. We compute the segmentation masks as:
To generate binary segmentation masks, we take operation over the K at every pixel location and encode the maximum indicator (indices) using one-hot codes. We render a scene image using a composition of all scene objects as:
D. Additional Results
To compare quantitatively the intra-object disentanglement achieved by MulMON and IODINE, we employ the framework and metrics (DCI) of Eastwood and Williams. Specifically, let be the ground-truth generative factors for a single-view of a single object in a single scene, and let be the corresponding learned representation. Following , we learn a mapping from to with random forests in order to quantify the disentanglement, completeness and informativeness of the learned object representations. Section 5.3 presented the results on the CLE-MV dataset, and here we present the results on the CLE-Aug dataset. As shown in Table 6, MulMON again outperforms IODINE, learning representations that are more disentangled, complete and informative (about ground-truth factor values). It is worth noting the significant gap in informativeness () in Table 6. This strongly indicates that the object representations learned by MulMON are more accurate, i.e. they better-capture object properties.
D.2 Generalization Results
To evaluate MulMON’s generalization ability, we trained MulMON, IODINE and GQN on CLE-Aug. Then, we compared their performance on CLE-MV and 2 new datasets—Black-Aug and UnseenShape (see Figure 9). Black-Aug contains the CLE-Aug objects but only in single, unseen colour (black). This tests the models’ ability perform segmentation without colour differences/cues. UnseenShape contains only novel objects that are not in the CLE-Aug dataset—cups, cars, spheres, and diamonds. This directly tests generalization capabilities. Both datasets contain 30 scenes, each with 10 views.
Table 1 shows that 1) all models generalize well to novel scenes, 2) MulMON still performs best for all tasks but observation prediction—where GQN does slightly better due to its more direct prediction procedure (features layout vs. features objects layout), 3) MulMON can indeed understand the composition of novel objects in novel scenes—impressive novel-view predictions (observations and segmentations) and disentanglement. Despite the excellent quantitative performance achieved by MulMON in generalization, we discovered that MulMON tended to decompose some objects, e.g. cars, into pieces (see Figure 10). Future investigations are thus needed in order to enable MulMON to generalize to more complex objects.
D.3 Ablation Study
Prediction performance vs. the number of observations T
In Section 5.4 of the main paper, we show that the spatial uncertainty MulMON learns decrease as more observations () are acquired. Here we study the effect of on MulMON’s task performance, i.e. novel-viewpoint prediction. We employ mIoU (mean intersection-over-union) and RMSE (root-mean-square error) to measure MulMON’s performance on observation prediction and segmentation prediction respectively. In Figure 11, we show that the spatial uncertainty reduction (Left) suggests boosts of task performance (Right). This means MulMON does leverage multi-view exploration to learn more accurate scene representations (than a single-view configuration), this also explains the performance gap (see Section 5.1 in the main paper) between IODINE and MulMON. To further demonstrate the advantage that MulMON has over both IODINE and GQN, we compare their performance in terms of both segmentation and novel-view appearance prediction, as a function of the number of observations given to the models. Figure 12 shows that; 1) MulMON significantly outperforms IODINE even with a single view, likely due to a superior 3D scene understanding gained during training (figures on the left), 2) Despite the more difficult task of achieving object-level segmentation, MulMON closely mirrors the performance of GQN in predicting the appearance of the scene from unobserved viewpoints (figures on the right), 3) MulMON achieves similar performance in scene segmentation from observed and unobserved viewpoints, with any difference diminishes as the number of views increase (see dashed lines vs. solid lines in the left-hand figures).
Although explicit assumption about the number of objects in a scene is not required for MulMON, selecting a appropriate (i.e. the number of object slots) is crucial to have MulMON work correctly. In the main paper, we discussed that “ needs to be sufficiently larger than the number of scene objects” and we show the experimental support here. We train our model on CLE-MV, where each scene contains 4 to 7 objects including the background, using and run tests on novel-viewpoint prediction tasks using various . Figure 13 shows that, for both observation prediction and segmentation prediction tasks, the model’s performance improves as increases until reaching a critical point at , which is the maximum quantity of scene objects in the dataset. Therefore, one should select a that is always greater or equal to the maximum number of objects in a scene. When this condition is satisfied, further increase will mostly not affect MulMON’s performance.
Subtle cases in terms of ’s selection do exist. As shown in Figure, instead of treating the Shep7 scene a combination of a single object and the background, MulMON performs as a part segmenter that discovers the small cubes and their spatial composition. This is because, in the training phase of MulMON, the amortized parameters and are trained to capture the object features (possibly disentangled) shared across all the objects in the whole dataset instead of each scene with specific objects. These shared object features are what drives the segmentation process of MulMON. In Shep7, what is being shared are the cubes, the object itself is a spatial composition of the cubes. The results on Shep7 along with the results shown in Figure 10 illustrate the subjectiveness of perceptual grouping and leaves much space for us to study in the future.
Effect of the IG coefficient in scene learning
In the MulMON ELBO, we fix the coefficient of the information gain at 1. In testing, we consider this coefficient controls the scene learning rate (uncertainty reduction rate). We denote the coefficient as hereafter. According to the MulMON ELBO (to maximize), the negative sign of the term suggests that greater value of the coefficient leads to less information gain (spatial exploration). To verify this, we try four different (0.1, 1.0, 10.0 and 100.0) and track the prediction uncertainty as observations are acquired (same as our ablation study of ). The results in Figure 15 verifies our assumption the scene learning rate: larger leads to slower scene learning and vice versa.
D.4 Random Scene Generation
As a generative model, MulMON can generate random scenes by composing independently-sampled objects. However, to focus on forming accurate, disentangled representations of multi-object scenes, we must assume objects are i.i.d. and thus ignore inter-object correlations—e.g. two objects can appear at the same location. Figure 16 shows some random scene examples generated by MulMON (trained on the CLE-MV dataset). We can see that MulMON generates mostly good object samples by randomly composing different features but does not take into account the objects’ spatial correlations.