Planar Surface Reconstruction from Sparse Views
Linyi Jin, Shengyi Qian, Andrew Owens, David F. Fouhey
Introduction
Consider the two photos in Figure LABEL:fig:teaser, as humans, we can infer that they were taken from the same scene: there is a chair on one side, a bedside table on the other, and a large glass wall and floor in the middle. We perceive the scene correctly despite the fact that they are sparse views : they were taken from very different camera poses, and very little of the scene structure within them overlaps. There are also challenges in grouping: while each photo might, at first glance, seem to contain its own glass wall and floor, they are each “slices” of the same planar objects. Despite these challenges, humans readily understand spaces like these from only a few ordinary photos, such as when they share collections of photos from the same event or look for housing.
Yet this setting poses challenges for today’s computer vision methods. Traditional tools from multi-view geometry largely rely on correspondence for reconstruction and are fundamentally limited to the small part of the scene that directly overlaps, even when the camera pose is known. Learning-based single view 3D , offers ways of reconstructing each image, but produces two messy piles of partial reconstructions due to the unknown viewpoint change. In Figure LABEL:fig:teaser, the floor is fragmented across both views and the back of the chair is present in only one view. While humans can associate these pieces and infer their relative positions to produce a coherent reconstruction, it is not trivial for today’s reconstruction algorithms.
We believe the ease with which humans solve the two unknown camera reconstruction problem, coupled with the difficulty it poses for computers marks it as an important task on the path to human-level 3D perception. Indeed, it poses challenges to existing work in deep learning for multiview reconstruction, which typically requires known camera poses as opposed to unknown poses, many views as opposed to two, additional depth information at test time as opposed to RGB images, or works only on synthetic data . Typically the extra information used is fundamental to the algorithm (e.g., using poses for triangulation) and cannot be removed to produce a method that works in the two unknown view reconstruction settings.
We propose a learning-based approach that constructs a coherent 3D reconstruction from two views with an unknown relationship. Our insight, supported empirically, is that progress can be made by jointly tackling three related challenges: per-view reconstruction, inter-view correspondence, and inter-view 6DOF (rotation and translation) pose. Throughout, we use plane segments as our representation since they have simple parameters, are often good approximations , and there is a strong line of work for estimating them or related properties from images.
Our method, described in Section 3, combines a deep neural network architecture and an optimization problem to jointly estimate planes, their relationships, and camera transformations. Our architecture builds on PlaneRCNN to produce, per-input, plane segments and parameters, per-plane embeddings for correspondences, as well as a probability distribution over relative cameras. This information is used in a discrete-continuous optimization problem (along with optional point features) to produce a coherent reconstruction across views. This reasoning across views enables our approach to, for instance in Figure LABEL:fig:teaser, produce a single floor rather than a set of inconsistent floor fragments, and jointly infer the distance from the images to the scene boundaries as well as the relative camera pose.
We validate our approach on realistic renderings from the Matterport3D dataset using pairs with limited overlap (average rotation, m translation, overlap). We report experimental results in Section 4 for three tasks: producing a single coherent reconstruction from the two views, matching planes across views, and estimating the full 6DOF relative camera pose. We compare extensively with a variety of baselines (e.g., adding an independent network to estimate relative pose followed by fusion of the two scene layouts) and ablations that test the contributions of our method and design choices. Our results demonstrate the value of joint consideration of the interrelated problems: our approach substantially outperforms the fusion of existing approaches to the independent problems of camera pose estimation and scene structure estimation.
Related Work
The goal of this paper is to produce a coherent 3D plane surface reconstruction of the scene given two images with an unknown relationship between the cameras. To solve the problem, our approach needs to both reconstruct 3D shapes from 2D images and establish correspondence and identify the relative camera pose between views. Our work therefore touches on many topics in 3D computer vision, ranging from single-view reconstruction to correspondence to relative camera view prediction to two-view stereo.
Much of our signal comes from reconstruction methods that map 2D images to 3D structure. This has long been a goal of computer vision with methods that aim to extract normals , voxels , and depth from 2D images. Our approach builds most heavily on a work aiming to produce a planar reconstruction . We build upon the work in this area, in particular PlaneRCNN , but we use it to build a planar reconstruction from two perspective images (i.e., what an ordinary person’s cell phone might capture casually). This focus on perspective images separates our work from approaches that use panorama images .
In the process, we find correspondence between planar regions of the scene. Correspondence is, of course, one of the long-standing problems in computer vision. A great deal of work aims to describe patch-based regions, ranging from classic SIFT descriptors to learned descriptors , which are often paired with projective geometry ; other work finds matches across object-level correspondence or planar-level correspondence . We see our work of jointly reconstructing and identifying correspondence as complementary to this work; a core component is an embedding network that aims to describe the plane, like and the last component of our system uses VIP-like features to constrain plane parameters. Our approach, however, solves a superset of these problems, since it produces a reconstruction as well.
We additionally predict the relative transformation between the cameras. This has been studied as an output of correspondence methods from both a classical and learning-based perspective. There are, however, methods that directly try to predict these transformations, including by deep networks , aligning RGBD data , or learning to predict the fundamental matrix . Like these approaches, we aim to estimate the relative transformation between the images as a component (although unlike some, we assume only ordinary RGB images); improving this component is complementary to our goals.
The relationship between reconstruction, correspondence, and camera pose, and the value of planes for inference has long been well-understood in the stereo community. While we use two views, these are well-separated and unknown, so our approach is different compared to standard two-view stereo or visual SLAM and more similar to wide-baseline stereo . Unlike wide-baseline stereo systems, however, we also produce reconstructions for portions of the scene that are seen in only one camera. Nonetheless, we draw inspiration from works in this community that use planes as a useful unit of inference .
The most similar work in this direction is Associative3D , which solves a related problem of reconstructing volumetric objects. Our approach is inspired by , but overcomes several methodological limitations, which we experimentally show. Associative3D uses one network for detecting objects and another for relative camera pose, and fuses the two via a heuristic RANSAC-like scheme. While effective on a six-object subset of the clean synthetic SUNCG dataset , these components fall short on the realistic Matterport3D dataset that our approach uses. Meeting the challenge of handling with planar segments (which cover objects and layouts) in realistic data requires a more principled optimization strategy and better signal for relative camera estimation. As another benefit, our approach has one backbone network forward pass per image.
Approach
Our approach, depicted in Figure 1, aims to map two images with an unknown relationship to a set of globally consistent planes and relative camera poses. This task requires estimating plane parameters, reasoning about their relationship (to avoid, e.g., a reconstruction with two floors), and inferring the relative pose. These subproblems are related since, for instance, plane parameters and correspondence constrain relative camera pose. We approach this with a network that predicts parameters and embeddings for planes (Sec. 3.1) and a distribution over relative camera pose (Sec. 3.2), followed by a joint optimization (Sec. 3.3) to produce a final coherent scene via joint reasoning. At training time, our system depends on RGBD for supervision, but can run on ordinary RGB images at test time.
Our plane prediction module produces, per-image, a set of plane segments that also have an embedding for cross-view matching. Each plane has: a segment ; plane parameters \mbox{\boldmath\pi}_{i}=[\mathbf{n}_{i},o_{i}] (where is a unit vector normal and the offset) giving the plane equation \mbox{\boldmath\pi}_{i}^{T}[x,y,z,-1]=0; and a unit-norm embedding for cross-view matching. Throughout, we denote planes in view 2 as \mathcal{M}_{j}^{\prime},\mbox{\boldmath\pi}_{j}^{\prime},\mathbf{e}_{j}^{\prime}, etc.
Plane detection. We adopt PlaneRCNN to detect planes in each image. During inference, the system produces backbone features using ResNet50-FPN . We use a region proposal network to propose boxes and then infer plane masks and normals from features from RoIAlign . At the same time, a decoder maps the backbone feature to a depthmap, which is used to compute the plane offset.
Appearance embedding. We additionally predict a cross-view embedding by training the network on pairs of images via triplet loss. Given plane correspondence, we form cross-view triplets where anchor corresponds with the positive match and not the negative match . We minimize a standard triplet loss , which gives a loss if the anchor and positive are not closer than the anchor and the negative by a margin . We use online triplet mining and randomly pick negative matches with positive loss.
2 Camera Pose Module
Our camera pose module estimates a distribution over the relative camera pose between views. This enables joint reasoning between a holistic estimate of the camera as well as the valuable cues in the estimated geometry: plane correspondences and parameters can constrain relative poses. We produce a distribution , over a set of discrete pairs of rotations and translations ).
Our network outputs two independent multinomial distributions over translation and rotation using attention-style features. The joint translation/rotation distribution is then their product. Our attention features follow recent literature that find similarity between images using attention and capture, at each pixel in a feature map, the relative similarity between that pixel and the pixels in the other image. We use the backbone network to extract -dimensional feature maps of size and . Suppose and index pixels in feature and over both rows and columns (i.e., is a -dimensional column), we then compute the attended feature :
We reshape this to a -channel feature map where each pixel represents the normalized correlation with each of the other pixels in feature map 2. We apply a convolutional network to the reshaped with six layers of convolutions and two fully connected layers that predicts the camera distribution. We found this attention to be superior compared to the strategy of concatenating average-pooled vector outputs in . as well as other alternate strategies in experiments in Section 4. We train the camera pose module on top of the ResNet50-FPN’s feature .
3 Optimization
Our final step is an optimization that serves two purposes. First, it propagates information between the related problems: the parameters of a mutually visible plane can inform relative camera poses, and a better pose of a camera can help disambiguate matches. Second, it ensures the coherence: naively concatenating two single-view parses of the scene (as we show empirically) yields a collection of possibly overlapping, inconsistent planar fragments showing the same object from different vantage points.
We cast this as an optimization problem over camera-to-camera transformation and set of plane parameters \hat{\mbox{\boldmath\pi}}_{i},\hat{\mbox{\boldmath\pi}}_{j}^{\prime} as well as a plane correspondence matrix where is if and only if plane corresponds to plane . As input, the optimization has a series of planes for the view 1 \{\mathcal{M}_{i},\mbox{\boldmath\pi}_{i},\mathbf{e}_{i}\} and planes for view 2 \{\mathcal{M}_{j}^{\prime},\mbox{\boldmath\pi}_{j}^{\prime},\mathbf{e}_{j}^{\prime}\}, and a distribution over relative camera transformations . The optimization tightly couples discrete and continuous variables and has many degenerate solutions. We thus follow a two stage approach: we hold plane parameters fixed and solve for plane correspondence and rough camera location; we then continuously optimize camera and plane parameters.
Discrete Problem: We first solve a discrete problem that selects a camera from the options and the plane correspondence matrix . The most important term is expressed via a cost matrix encoding the quality of a plane correspondence assuming camera has been selected. When the second view’s plane parameters have been transformed into the first view by camera (omitted for clarity), the cost matrix (with trade-off parameters ) is
which includes terms for the embedding, normal, and offset which are with perfect matches.
This cost term is combined with two regularizing terms. One, , penalizes unlikely cameras via the negative log-likelihood of camera ; the other rewards matching objects: . We pick the best correspondence and camera that optimize the objective
which encourages selecting a likely camera, as many planes as possible, and correspondence that is consistent in both appearance and geometry assuming the camera is correct.
With a fixed camera , the objective can be efficiently solved using the Hungarian algorithm, and the number of matches handled via thresholding. Since there are a finite number () of camera hypotheses, this amounts to solving independent matching problems.
Continuous Problem: Having selected a camera and planar correspondences , we can then refine the predictions from the deep networks for the camera and planes. We optimize over camera transformations and plane parameters \hat{\mbox{\boldmath\pi}}_{i},\hat{\mbox{\boldmath\pi}}_{j}^{\prime} to minimize geometric distance between the corresponding planes (assuming the same coordinate frame) and pixel alignment error based on pixel-level features:
where measures the Euclidean distance of back-projected points that match across corresponding planes. Inspired by , we warp texture to viewpoint normalized cameras using the plane parameters to extract viewpoint-invariant SIFT features. The term regulates the deviation from the selected camera bin. We initialize the optimization at \mathbf{R}_{\hat{k}},\mathbf{t}_{\hat{k}},\mbox{\boldmath\pi}_{i},\mbox{\boldmath\pi}_{j}^{\prime}, and optimize with a trust-region reflective minimizer. We parameterize rotations as 6D vectors following .
Merging Planes: Given planes in correspondence and camera transformations, we merge them in the global frame. We merge offsets by averaging and normals by solving for the maximizing via an eigenvalue problem.
4 Implementation Details
A full detailed description appears in the supplemental. We use Detectron2 to implement our network. The plane detection backbone uses ResNet50-FPN pretrained on COCO . We directly regress normals instead of using classification in . The camera branch and the plane embedding are trained using a Siamese network whose backbone is the plane detection backbone. During training, we first train the plane prediction backbone on single images and then freeze the network. We train the plane appearance embedding and camera pose module on the frozen backbone. We fit all trade-off parameters (e.g., ) on the validation set using randomized search. We run -means clustering and spherical -means clustering on the training set to produce 32 bins for translation and rotation respectively.
Experiments
We evaluate our approach using renderings of real-world scenes. Our approach generates a new, rich and coherent output in terms of a planar reconstruction plus a geometric relationship between the two views. In the process of producing this reconstruction, it solves two other problems: predicting object correspondences and relative camera pose estimation. We therefore evaluate our model in three ways, which each naturally have their own metrics and baselines: the full model (Section 4.2), the plane correspondence (Section 4.3) and the relative camera poses (Section 4.4).
We use renderings of real scenes from Habitat and Matterport3D . Habitat enables the rendering of realistic images and provides ground truth pose and depth that can be used for evaluation. We stress that our image pairs have far less overlap compared to other settings with an average overlap of just vs used by DeMoN on Sun3D or 80% by the plane odometry method (see supplement for more statistics).
Dataset: Our training, validation, and test set consist of 31932, 4707, and 7996 image pairs respectively. We fit our ground truth planes on Matterport3D using a standard RANSAC approach, using semantic labels to constrain the planes, following . We fit planes on the 3D meshes and render them to 2D, which provides us with plane correspondence even when the planes have limited overlap. We generate camera poses following and select pairs with common planes and unique planes like .
2 Full Scene Evaluation
We begin by evaluating our approach’s ability to produce full scene reconstruction in terms of a set of 3D plane segments in the scene in a common coordinate frame.
Metrics: We follow other approaches that reconstruct the scene factored into components and treat the full problem like a detection problem, evaluated using average precision (AP). Given a ground-truth decomposition into components (in this case, planes), we evaluate how well the approach detects and successfully reconstructs each one. We define a true positive as a detection satisfying three criteria: (i) (Mask) mask intersection-over-union ; (ii) (Normal) surface normal distance, ; and (iii) (Offset) offset distance, m.
Baselines: We compare with alternate approaches and ablations that test the system’s contributions. To the best of our knowledge, no existing work solves our task, so we test fusions of existing systems. These baselines estimate a relative camera pose and per-image planes.
Plane Odometry + MWS / PlaneRCNN : extracts plane-primitives from successive RGB-D video frames and uses plane-to-plane registration to solve for the camera pose. For the 3D representation, we either predict planes using PlaneRCNN in our plane prediction module or run Manhattan-world Stereo on predicted depth (MWS) or ground truth depth (MWS-G). To provide and with predicted depthmaps, we use MiDaS , a state of the art system trained on M images. We found that MWS gave better plane fits on depthmaps without semantic labels compared to our RANSAC approach, also shown in . Plane Odometry mainly reasons with geometric information when producing its final output and tests a simple fusion of visual odometry and plane extraction systems.
SuperGlue-G + MWS / PlaneRCNN : is a point-based matching algorithm that uses deep features to find pixel correspondence across images to estimate an essential matrix. The scale of the translation is not determined by the essential matrix so we use the ground truth translation scale for this method. Approaches like are complementary to our camera branch since they have finer resolution but no object-size prior; our approach could integrate in its continuous optimization. Like the previous baseline, we try each of MWS/MWS-G/Plane-RCNN.
RPNet + PlaneRCNN : We apply the PlaneRCNN module used by our system and our improved variant of RPNet – see Sec. 4.4 for more details. This is equivalent to our system, but ignoring our optimization or reasoning over plane parameters. This tests the value of joint reasoning while holding base networks fixed.
Appearance Embedding Only: We run the optimization with only the appearance embedding predicted for each plane. This gives a fusion of planes based on learned embeddings, which tests the value of our geometric reasoning.
Associative3D Optimization: We apply a heuristic RANSAC-style optimization scheme from the most closely related paper . We adapt it to our setting for a fair comparison (see supplemental for detail). This tests the value of our optimization, which casts camera selection and matching in terms that are solvable via principled algorithms.
No Continuous Optimization: We perform the full method without the continuous optimization in Section 3.3.
Proposed: We apply the full proposed method.
Qualitative Results: We show qualitative results on Matterport3D test set in Figure 2. Prediction and ground truth are shown in two novel views to see all planes and camera poses in the scene. Our method predicts accurate camera poses from inputs with small overlap, merges planes across different views, and aligns their textures to produce a coherent reconstruction. More examples are in the supplement.
We show comparison between our proposed approach and baselines in Figure 3. Plane Odometry fails on predicted depth due to distortion and rarely finds the correct correspondences between planes, e.g. incorrect walls are aligned. SuperGlue-G can usually obtain a reasonable camera pose given oracle translation scale but does not merge planes across views, e.g., non-parallel walls. No Continuous Optimization finds and merges the corresponding planes using the Hungarian algorithm in our discrete optimization. However, simply merging the corresponding planes still generates erroneous results, e.g., unaligned texture of pictures on the wall in row 1, disconnected walls in row 2, which are caused by discretization bins of the camera transformation and inaccurate single-view prediction. Our proposed approach refines the camera and planes further, so that it fixes the inconsistency of plane detection and camera modules and generates reconstructions that are significantly closer to the ground truth.
Quantitative Results: We show results in Table 1. Fitting planes on predicted depthmaps (MWS) fails poorly due to distorted predictions. This is a known challenge (e.g., addressed by concurrent work without public code at time of submission). Methods using PlaneRCNN outperform MWS even with ground truth (GT) depth, also found in ; this is because of PlaneRCNN’s stronger plane labels improved by semantics. Plane Odometry finds pose primarily with plane normals. Its correspondence is often wrong because our low-overlap setting limits the information in geometric cues alone (in contrast to the combined normal, offset, and texture used by our method). SuperGlue with GT translation scale (SuperGlue-G) slightly outperforms the RPNet with PlaneRCNN, but falls short of the proposed full method by 6 AP. Some of this gap is because there are duplicate copies of planes found in both images. One contribution of our method is preventing this case. While merging based on appearance only partially solves the problem, the full method, which incorporates geometry, does the best. The particular optimization for this merging is moreover important: the heuristic randomized optimization of does worse. The continuous optimization makes little difference in planes but improves the camera considerably (Sec. 4.4). Additional experiments on AP vs. overlap are in the supplement.
3 Plane Correspondence
One core challenge for generating a single coherent reconstruction from two images is ensuring that each visible plane appears exactly once rather than being split. We therefore evaluate how well we can identify correspondences between planes across views.
Baselines: We compare the proposed approach with three baselines that test the contribution of each component of our method. We measure performance using ground-truth bounding boxes during evaluation so that we measure only errors due to correspondence and not detection.
Appearance Only: We apply the Hungarian algorithm to the appearance embedding and do not use geometry. This tests how important geometry is to identifying correspondence.
MessyTable ASNet on ROI: We adapt the ASNet framework to our setting and architecture for a fair comparison (see supplemental for details). This network improves matching by adding an additional pathway that has more context. This tests whether our matching is constrained by looking only inside a proposal region.
Associative3D Optimization: We again apply the optimization method of , following Section 4.2.
Metrics: We aim to measure how well methods associate planes in a pair of images. We therefore evaluate performance directly with Image Pair Association Accuracy (IPAA) metric by Cai et al. , which represents the fraction of image pairs with no less than X% of the planes associated correctly (written as IPAA-X). We use ground truth bounding boxes in this section for evaluation.
Quantitative Results: Table 2 shows that our method outperforms all other benchmarks on IPAA-100 and IPAA-90 and has a slight drop of 0.4% on IPAA-80. Our full optimization improves IPAA-100 by a large margin (9.4%) compared to using the appearance embedding matrix only. We hypothesize our approach outperforms due to the factorial growth of the search space, which the heuristic approach struggles with. We suspect that struggles since it is designed for table-top matching, where the planar table constrains how much context can change from view to view (in contrast to our strongly non-planar scenes).
Qualitative Results: Appearance Only and ASNet do not consider geometric information; therefore, they have difficulty distinguishing planes of similar texture, e.g., two sides of the bed in Figure 4. Associative3D does not find all the correspondence due to its random search scheme. Our method detects all the correspondence and distinguishes similar texture planes using geometric information.
4 Relative Camera Pose Estimation
Our final experiment tests the effectiveness of the predicted relative camera pose. We compare our method with a number of alternate approaches and baselines. Additional experiments testing the use of attention features rather than concatenation features appear in the supplement.
Metrics: Following prior works , we measure the camera rotation and translation by geodesic rotation distance and Euclidean distance separately. We report the mean and median error, as well as the fraction below a fixed threshold ( and m following ). Evaluation on the translation angle error is in the supplement.
Baselines: We compare our full system with four baselines and two ablations.
Plane Odometry + GT Depth: We apply the RGBD plane odometry system from to the ground-truth depth. As described in Section 4.2, this approach mainly reasons geometrically, without learned priors or deep features.
Plane Odometry + : We apply the same approach, using estimated depth from for a more fair comparison.
Associative3D camera branch: We apply an improved version of the camera branch of Associative3D, which is an improved version of the RPNet . This method average pools ResetNet50 features followed by a MLP. For a fair comparison, we upgrade the MLP to match the number of layers and parameters of our approach. This tests the value of our improved attention features.
: We apply the SuperGlue system. Although as a pixel-correspondence-based method, we see this as a complementary line of work to our method. SuperGlue estimates camera poses via the essential matrix and thus cannot estimate the translation’s scale. We note that the approach in Section 4.2 used the ground-truth camera translation scale to give an upper bound if the true translation scale were used. The approach fails on of our pairs due to insufficient correspondence; in this case, we assign the identity matrix for rotation. We use the only available model, which was trained on ScanNet .
No Optimization: We apply our camera branch and take the most likely camera without any joint reasoning.
No Continuous: We apply our approach but do not apply continuous optimization.
Quantitative Results: We report results in Table 3. The full method has low median errors of m/. Comparing with other approaches, we find that our approach also substantially outperforms the Plane Odometry method regardless of whether ground truth or predicted depth is used. Just as described in Section 4.2, this is due to the limited amount of information that is available in geometry alone. SuperGlue can be extraordinarily accurate in rotation when textures provide strong correspondences across views, but also can fail to make inferences on hard scenes.
The most likely camera from our approach (No Optimization) substantially outperforms the previous method of , which follows a simpler average pooling. This is largely due to our use of attention features; additional more detailed ablations are in the supplement. Discrete optimization consistently slightly improves results, but continuous optimization substantially improves translation with marginal gains and loses in rotation performance.
Conclusion
We’ve presented a learning-based system to produce a coherent planar surface reconstruction from two unknown views. Our results suggest that jointly considering correspondence and reconstruction improves both, although our experiments suggest room for growth (e.g., see typical failure modes in Figure 5, including low-overlap planes and repeated coplanar and similar segments). Future directions include using our framework to reconstruct a more complete scene with fewer frames than traditional SfM . Figure 6 shows examples that extend our system to views by incrementally stitching new views on reconstructed views.
Acknowledgments We thank Dandan Shan, Mohamed El Banani, Nilesh Kulkarni, Richard Higgins for helpful discussions. Toyota Research Institute (“TRI”) provided funds to assist the authors with their research but this article solely reflects the opinions and conclusions of its authors and not TRI or any other Toyota entity.
References
Appendix A Additional Results
To give additional context for the dataset, we transform the point cloud of view 1 to view 2. The average percentage of overlapping points of our dataset is . For context, Sun3D image pairs used by DeMoN had an average overlap of ; FR2_desk and FR3_structure used by Plane Odometry had an overlap of . Figure 8 shows the histogram of overlap ratio of our dataset versus the other datasets.
Random examples are shown in Figure 9. Our dataset is much more sparse than datasets used in DeMoN and Plane Odometry, because they sample adjacent frames from a video, while we render a sparse view dataset ourselves.
A.2 Performance vs. Overlapping Ratio
Figure 7 shows the performance of our method and other baselines with images of different overlap ratios. Our method has higher AP over other baselines across all overlap ratios in our dataset. Overlap ratio of around 0.05 is an extremely hard setting, our method (29 AP) and improved RPNet (27 AP) perform better than matching based methods (24 AP) because SuperGlue rarely finds enough correspondence. Deep networks, on the other hand, can use other cues and prior to make predictions. When the overlap ratio is higher, our method can also outperform other baselines (increase by over 5 AP when overlap ratio is around 0.25 and 11 AP when overlap is about 0.55). We can merge planes based on correspondence to produce a coherent reconstruction while other baselines produce inconsistent and duplicated planes.
A.3 Rank by Single-view Prediction
We further investigate how the quality of single-view prediction (outputs from plane prediction module on each image) will affect our results. Table 4 shows results on top 25%, 50%, 75% and 100% of the test examples ranked by single-view AP. On top 25% examples where the single-view prediction are more accurate, our proposed methods have a higher AP (increase by over 10 points compared to 100% data). Moreover, our proposed method has a higher gain over ablation baselines Associative3D optimization and Appearance Embedding only by above 4.9 points on top 25% examples and even higher ( points) over the other external baselines. It is still significant but slightly lower when we have worse single-view predictions.
A.4 Translation Direction Evaluation
In additional to Table 3, we report a scale-free angle between the predicted and GT vectors in Table 5, following SuperGlue . When fails (11.6% of the data), we use a vertical vector (equidistant to all horizontal directions). The results are similar to rotation: does better if it has good correspondences, but can fail entirely. So its median error is lower, but other errors are comparable. The ablation shows our continuous optimization improves on using discrete cameras. We stress that our system recovers scale, which is not evaluated here, and that we see pixel-based approaches as complementary – one could also train a SuperPoint-like system on early layers in our network and use it in the optimization.
A.5 Ablation Study for Camera Pose Module
We include detailed architectures for baselines in Section 4.4 of our paper in Table 9, 10 and 11. ReLU is used between all Linear and Conv layers. We use ResNet50 pretrained on COCO as the backbone to predict camera pose. We do not freeze the backbone for the baselines since they are standalone networks. Table 6 shows the ablation experiments. With the ability to explicitly calculate the relationship of features across views, our proposed attention module (ResNet50-Attention) outperforms other standalone architectures by a large margin on all the metrics. Running the attention module on the plane detection backbone further improves results, but due to switching to a FPN , it is difficult to directly ascribe performance changes.
A.6 Qualitative Results
We first present our reconstruction results on selected examples in Figure 11, extending Figure 2 in our paper. We show our prediction and ground truth from two novel views to see all planes in the whole scene – a slightly raised view and a top down view. We then present results automatically evenly spaced in the test set in Figure 12, according to single-view AP. As we quantitatively show in Section A.3, higher single-view AP means much better results. Therefore, we hope the evenly spaced results can represent the overall performance of our approach better. When single-view plane prediction is accurate, our reconstructions from sparse views are typically reasonable.
We also show more selected qualitative examples on plane correspondence prediction in Figure 13, extending Figure 4 in our paper. Those examples use predicted bounding boxes. To determine the ground truth correspondence, we assign each ground truth box with a predicted box whose mask IoU is greater than 0.5. We also randomly choose examples in Figure 14 to reflect the overall performance of our approach and baselines. Those examples use ground truth boxes which are the same as the ones used in correspondence evaluation. As shown in the random results, there are many false positives and false negatives in challenging cases; therefore, there is still much space to improve in predicting plane correspondence.
A.7 Generalization to other data.
To test generalization, we also test our approach on images from ScanNet and Replica . Figure 10 shows results from our model trained on Matterport3D. Our model can obtain a reasonable interpretation.
Appendix B Implementation
We augment Matterport3D with plane segmentation annotations on the original mesh. This enables consistent plane instance ID and plane parameters through rendering, and automatically establishes plane correspondences across different images. We fit planes using RANSAC, following , on mesh vertices within the same object instance (using the annotation provided by the original Matterport3D dataset). We then render the per-pixel plane instance ID, along with the RGB-D images using AI Habitat . Since the original instance segmentation mesh has “ghost” objects and holes due to artifacts, we further filter out bad plane annotations by comparing the depth information and the plane parameters. Figure 15 shows examples of plane annotations on the mesh and rendered plane segmentations. Small plane masks with area less than 1% of the total image pixels are removed.
We follow to generate random camera poses. To include more data, we keep all valid cameras in each horizontal sector. The camera is of a random height 1.5-1.6m above the floor and a downward tilt angle of 11 degrees to simulate human’s view. The same camera intrinsic are used to render all the images. To generate image pairs, we randomly sample cameras within each room, and enforce that there are at least three common planes and at least three unique planes in each image. Our ground truth data includes a depthmap and plane segmentation for each image, as well as plane correspondences and camera transformations for the image pair.
B.2 Network Architectures
We use Detectron2 to implement our network. The overall architecture of our network is shown in Table 7. The backbone, RPN, box branch and mask branch are identical to Mask R-CNN . The depth branch is the same as the depthmap decoder in . Table 8 shows the exact architecture of the camera pose module in our model.
B.3 Optimization
Associative3D optimization. We re-build the heuristic optimization on the same network branches our approach uses to ensure apples-to-apples comparisons; we also optimize their search to accommodate for differences in plane-vs-object matching by using Eqn. 3 as the objective.
Discrete optimization. Our discrete optimization searches through all camera options, (), and outputs the best camera as well as a binary correspondence matrix for planes across images using the Hungarian algorithm, shown in Algorithm 1. cam2world converts plane parameters from the camera frame to the world frame,
B.4 ASNet [3] Baseline.
Following the Appearance-Surrounding Network in , in addition to the appearance branch, we add a surrounding branch for feature extraction to augment our embedding head. We combine both appearance features and surrounding features to determine plane correspondences. Instead of cropping the original image which causes one inference per detection, we increase the receptive field of RoIAlign by a factor of 2 to extract larger ROI features. After RoIAlign, the features are cropped and passed to a surrounding extractor and an appearance extractor. Both the extractors use the same number and size of Conv layers and linear layers as our embedding head uses (4 Conv layers and 1 FC layer). We train using cosine similarity weighted loss as described in Equation 1 and 2 of . We ensure apples-to-apples comparisons by re-training the embedding head on the same freezed backbone our approach uses.