TransformerFusion: Monocular RGB Scene Reconstruction using Transformers
Aljaž Božič, Pablo Palafox, Justus Thies, Angela Dai, Matthias Nießner
Introduction
Monocular 3D reconstruction is a core task in 3D computer vision, aiming to reconstruct a complete and accurate 3D geometry of an object or an environment from only 2D observations captured by an RGB camera. A geometric understanding is key to applications such as robotic or autonomous vehicle navigation or interaction, as well as model creation and scene editing for augmented and virtual reality. In addition, geometric scene reconstructions form the basis for 3D scene understanding, supporting tasks such as 3D object detection, semantic, and instance segmentation .
While state-of-the-art SLAM systems achieve robust and scale-accurate camera tracking leveraging both visual and inertial measurements, dense and complete 3D reconstruction of large-scale environments from monocular video remains a very challenging problem – particularly for interactive settings. Simultaneously, notable progress has been made on multi-view depth estimation, estimating depth from pairs of images by averaging features extracted from the images in a feature cost volume . Unfortunately, averaging features across a full video sequence can lead to equal-weight treatment of each individual frame, despite some frames possibly containing less information in various regions (e.g., from motion blur, rolling shutter artifacts, very glancing or partial views of objects), making high-fidelity scene reconstruction challenging.
Inspired by the recent advances in natural language processing (NLP) that leverage transformer-based models for sequence to sequence modelling , we propose a transformer-based method that fuses a sequence of RGB input frames into a 3D representation of a scene at interactive rates. Key to our approach is a learned feature fusion of the video frames using a transformer-based architecture, which learns to attend to the most informative image features to reconstruct a local 3D region of the scene. A new observed RGB frame is encoded into a 2D feature map, and unprojected into a 3D volume, where our transformer learns a fused 3D feature for each location in the 3D volume from the image view features. This enables extraction of the most informative view features for each location in the 3D scene. The 3D features are fused in coarse-to-fine fashion, providing both improved reconstruction performance as well as interactive runtime. These features are then decoded into high-resolution scene geometry with an MLP-based surface occupancy prediction.
In summary, our main contributions to achieve robust and accurate scene reconstructions are:
Learned multi-view feature fusion in the temporal domain using a transformer network that attends to only the most informative features of the image views for reconstructing each location in a scene.
A coarse-to-fine hierarchy of our transformer-based feature fusion that enables an online reconstruction approach running at interactive frame-rates.
Related Work
Estimating depth from multi-view image observations has been long-studied in computer vision. COLMAP introduced a patch matching based approach which achieves impressive accuracy and remains established as one of the most popular methods for multi-view stereo. While COLMAP offers robust depth estimation for distinctive features in images, the patch matching struggles to densely reconstruct areas without many distinctive color features, such as floor and walls. Recently, learning-based approaches that build data-driven priors from large-scale datasets have improved depth estimation in these challenging scenarios. Some proposed methods rely only on a 2D network with multiple images concatenated as input . Several recent approaches instead build a shared 3D feature cost volume in reference camera space using feature averaging . These approaches estimate the reference frame’s depth within a local window of frames, but some also propagate information from previously estimated depth maps by using probabilistic filtering , a Gaussian process , or an LSTM bottleneck layer . Such multi-view depth estimation approaches predict single-view depth maps, which must be fused together to construct a geometric 3D representation of the observed scene.
D reconstruction from monocular RGB input.
Multi-view depth estimation approaches can be combined with depth fusion approaches, such as volumetric fusion , to obtain a volumetric reconstruction of the observed scene. MonoFusion is one of the first methods using depth estimate from a real-time variant of PatchMatch stereo . However, fusing noisy depth estimates causes artifacts in the 3D reconstruction, which lead to the development of recent approaches that directly predict the 3D surface reconstruction instead of per-frame depth estimates. One of the first approaches to predict 3D surface occupancy from two input RGB images is SurfaceNet , which converts volumetrically averaged colors into 3D surface occupancies using a 3D convolutional network. Atlas extends this approach to a multi-view setting, while also leveraging learned features instead of colors. Recently, NeuralRecon proposed a real-time 3D reconstruction framework, adding GRU units distributed in 3D to fuse reconstructions from different local windows of frames. Our approach also fuses together learned features from RGB frame input in an online fashion, but our transformer-based multi-view feature fusion enables relying only on the most informative features from the observed frames for a particular spatial location in the reconstructed scene, producing more accurate 3D reconstructions.
Transformers in computer vision.
The transformer architecture has achieved profound impact in many computer vision tasks in addition to its natural language processing origins. For a detailed survey, we refer the reader to . In computer vision, transformers have been leveraged successfully for tasks such as object detection , video classification , image classification , image generation , and human reconstruction . In this work, we propose transformer-based feature fusion for 3D scene reconstruction from a monocular video. Given a sequence of observed RGB frames, our approach learns to attend to the most informative features from each image to predict a dense occupancy field.
End-to-end 3D Reconstruction using Transformers
From these 2D image features, we construct a 3D feature grid in world space. To this end, we regularly sample grid points in 3D at a coarse resolution of every cm and a fine resolution of cm. For these coarse and fine sample points, we query corresponding 2D features in all images and predict fused coarse and fine 3D features using transformer networks :
Note that we also store the intermediate attention weights and of the first transformer layers for efficient view selection, which is explained in Sec. 3.4.
To further improve the features in the 3D spatial domain, we apply 3D convolutional networks and , at the coarse and fine level, respectively:
This extraction of surface occupancies is inspired by convolutional occupancy networks and IFNets . From this occupancy field we extract a surface mesh with Marching cubes . Note that in addition to surface occupancy, we also predict occupancy masks for near-surface locations at the coarse and fine levels. These masks are used for coarse-to-fine surface filtering (see Sec. 3.2), which improves reconstruction performance with a focus on the surface geometry prediction and enables interactive runtime.
We train our approach in end-to-end fashion by supervising the surface occupancy predictions using the following loss:
where and denote binary cross-entropy (BCE) losses on occupancy mask predictions for near-surface locations at the coarse and fine levels, respectively (see Sec. 3.2), and denotes a BCE loss for surface occupancy prediction (see Sec. 3.3).
As described above, denotes the attention values of the initial attention layer, which are used for view selection to speed-up fusion (see Sec. 3.4).
2 Spatial Feature Refinement
The refined features are also used to predict occupancy masks for near-surface locations at both coarse and fine levels, thus, filtering out free-space regions and sparsifying the volume, such that the higher-resolution and computationally expensive fine-scale surface extraction is performed only in regions close to the surface. To achieve this, additional 3D CNN layers and are applied to the refined features, outputting a near-surface mask for every grid point:
Only spatial regions where both and are larger than , i.e., close to the surface, are processed further to compute the final surface reconstruction; other regions are determined to be free space. This improves the overall reconstruction performance by focusing the capacity of the surface prediction network to close-to-the-surface regions and enables a significant runtime speed-up.
3 Surface Occupancy Prediction
We concatenate the interpolated features and predict the point’s occupancy as , where is a multi-layer perceptron (MLP) with 3 modules of feed-forward layers, containing ReLU activation, linear layer with residual connection, and layer norm.
4 View Selection for Online Scene Reconstruction
We aim to consider all frames as input to our transformer for each 3D location in a scene; however, this becomes extremely computationally expensive with long videos or large-scale scenes, which prohibits online scene reconstruction. Instead, we proceed with the reconstruction incrementally, processing every video frame one-by-one, while keeping only a small number of measurements for every 3D point. We visualize this online approach in Fig. 1.
During training, for efficiency, we use only random images for each training volume. At test time, we leverage the attention weights and of the initial transformer layers to determine which views to keep in the set of measurements. Specifically, for a new RGB frame, we extract its 2D features, and run feature fusion for every coarse and fine grid point inside the camera frustum. This returns the fused feature and also the attention weights over all currently accumulated input measurements. Whenever the maximum number of measurements is reached, a selection is made by dropping out a measurement with lowest attention weight before adding new measurements in the latest frame. This guarantees a low number of input measurements, speeding up fusion processing times considerably. Furthermore, by using coarse-to-fine filtering, described in Sec. 3.2, we can further accelerate fusion by only considering higher resolution points in the area near the estimated surface. Together with incremental processing that results in high performance benefits, our approach performs per-frame feature fusion at about 7 FPS despite an unoptimized implementation.
5 Training Scheme
Our approach has been implemented using the PyTorch library . The architecture details of the used networks are specified in the supplemental document. To train our approach we use ScanNet dataset , an RGB-D dataset of indoor apartments. We follow the established train-val-test split. For training, we randomly sample m volume chunks of the train scenes, sampling less chunks in free space and more samples in areas with non-structural objects, i.e. not only consisting of floor or walls. This results in k training chunks. For each chunk, we randomly sample RGB images among all frames that include the chunk in their camera frustums.
The 2D convolutional encoder for image feature extraction is implemented as a ResNet-18 network, pre-trained on ImageNet . During training, a batch size of chunks is used with an Adam optimizer with , , and weight regularization of . We use a learning rate of with k warm-up steps at initialization, and square root learning rate decay afterwards. When computing the losses of coarse and fine surface filtering predictions, a higher weight of is applied to near-surface voxels, to increase recall and improve overall robustness. Training takes about 30 hours using an Intel Xeon 6242R Processor and an Nvidia RTX 3090 GPU.
Experiments
To evaluate our monocular scene reconstruction, we use several measures of reconstruction performance. We evaluate geometric accuracy and completion, with accuracy measuring the average point-to-point error from predicted to ground truth vertices, completion measuring the error in the opposite direction, and chamfer as the average of accuracy and completion. To account for possibly different mesh resolutions among methods, we uniformly sample k points over mesh faces of every reconstructed mesh. Additionally, we threshold these point-to-point errors and compute precision and recall by computing the ratio of point-to-point matches within distance cm. Since it is easy to maximize either precision (by predicting only a few but accurate points) or recall (by over-completing reconstructions with noisy surface), we found the most reliable metric to be F-score, determined by both precision and recall.
Our ground truth reconstructions are obtained by automated 3D reconstruction from RGB-D videos of real-world environments and, thus, they are often incomplete due to unobserved and occluded regions in the scene. To avoid penalizing methods for reconstructing a more complete scene w.r.t. the available ground truth, we apply an additional occlusion mask at evaluation.
As most state of the art, particularly for depth estimation, rely on a pre-sampled set of keyframes (based on sufficient translation or rotation difference between camera poses), we evaluate all approaches based on sequences of sampled keyframes, using the keyframe selection of .
1 Comparison with State of the Art
In Tab. 2, we compare our approach with state-of-the-art methods. All methods are trained on the ScanNet dataset , using the official train/val/test split. We use the pre-trained models provided by the authors for MVDepthNet , GPMVS and DPSNet which are fine-tuned on ScanNet. For baselines that predict depth in a reference camera frame instead of directly reconstructing 3D surface, a volumetric fusion method is used to fuse different depth maps into a 3D truncated signed distance field. The single-view depth prediction method RevisitingSI suffers from the more challenging task formulation without the use of multiple views, leading to noisier depth predictions and inconsistencies between frames. Multi-view depth estimation methods leverage the additional view information for improved performance, with the LSTM-based approach of DeepVideoMVS achieving the best performance among these approaches. Reconstruction quality further improves with methods that directly predict the 3D surface geometry, such as NeuralRecon and Atlas . Our transformer-based feature fusion approach enables more robust reconstruction and outperforms all existing methods in both chamfer distance and F-score. The performance improvement can also be clearly seen in the qualitative comparisons in Fig. 3.
2 Ablations
To demonstrate the effectiveness of our design choices, we conducted a quantitative ablation study which is shown in Tab. 2 and discussed in the following.
We evaluate the effect of our learned feature fusion by replacing the transformer blocks with a multi-layer perceptron (MLP) that processes input image observations independently. The per-view outputs of this MLP are fused using an average (w/o TRSF, avg) or using a weighted average with weights predicted by the MLP (w/o TRSF, pred). We find that our transformer-based view fusion effectively learns to attend to the most informative views for a specific location, resulting in significantly improved performance over these averaging-based feature fusion alternatives.
Does spatial feature refinement help reconstruction performance?
Spatial feature refinement is indeed very important for reconstruction quality. It enables the model to aggregate feature information in spatial domain and produce more spatially consistent and complete reconstructions, without it (w/o spatial ref.) the geometry completion (and recall metric) are considerably worse.
How important is coarse-to-fine filtering?
Predicting the coarse and fine near-surface masks provides an additional performance improvement compared to the model without it (w/o C2F filter), as it allows more focus on surface geometry. Furthermore, this enables a speed-up of the fusion runtime by a factor of approximately , resulting in processing times of FPS (instead of FPS).
How many views should be used for feature fusion?
In our experiments, we use a limited number of frame observations to inform the feature for every 3D grid location. We find that these views all contribute, with performance degrading somewhat with sparser sets of observations ( or ). The number of frames is limited because of execution time and memory consumption for bigger scenes.
How effective is frame selection using attention weights?
The frames for each 3D grid feature are selected based on the computed attention weights and are updated during scanning. To evaluate this frame selection, we compare against a frame selection scheme that randomly selects frames that observe the 3D location (RND), which results in a noticeable drop in performance for both chamfer and F-score. The performance difference is even larger when using less views for fusion ( or ), where view selection becomes even more important. In Fig. 1, we visualize the most important view for locations in the scene, selected by the highest attention weight. Relatively smooth transitions between selected views among neighboring 3D locations suggest that view selection is spatially consistent. To illustrate the frame selection, we also visualize all selected frames with corresponding attention weights for specific 3D locations in the supplemental document.
3 Limitations
Under severe occlusions and partial observation of the scene, our method can struggle to reconstruct details of certain objects, such as chair legs, monitor stands, or books on the shelves. Furthermore, transparent objects, such as glass windows without frames, are often inaccurately reconstructed as empty space. We show qualitative examples of these failure cases in the supplemental material. These challenging scenarios are often not properly reconstructed even when using ground truth RGB-D data, and we believe that using self-supervised losses for monocular scene reconstruction could be an interesting future research direction. Additionally, higher resolution geometric fidelity could potentially be achieved by sparse operations in 3D or learning local geometric priors on detailed synthetic data .
Conclusion
We introduced TransformerFusion for monocular 3D scene reconstruction, leveraging a new transformer-based approach for online feature fusion from RGB input views. A coarse-to-fine formulation of our transformer-based feature fusion improves the effective reconstruction performance as well as the runtime. Our feature fusion learns to exploit the most informative image view features for geometric reconstruction, achieving state-of-the-art reconstruction performance. We believe that our interactive scanning approach provides exciting avenues for future research, and enables new possibilities in learning multi-view perception and 3D scene understanding.
Acknowledgments
This project is funded by the Bavarian State Ministry of Science and the Arts and coordinated by the Bavarian Research Institute for Digital Transformation (bidt), a TUM-IAS Rudolf Mößbauer Fellowship, the ERC Starting Grant Scan2CAD (804724), and the German Research Foundation (DFG) Grant Making Machine Learning on Static and Dynamic 3D Data Practical.
References
Appendix A Additional Results
In Fig. 4, we show a 3D reconstruction of a scene from the test set of the ScanNet dataset . To visualize the view selection approach presented in the main paper that is based on the attention weights of the used transformer networks, we render the camera views that are selected for a specific 3D location (green point) with corresponding attention weights (color temperature corresponding to the weight). On the right, we show the corresponding input images.
Additional Ablation Studies.
As analyzed in the main document, the different algorithmic parts of our methods play an important role. Fig. 5 shows qualitative results for the ablation study. Specifically, one can clearly see the impact of the spatial refinement as well as the temporal feature fusion via our transformer architecture. For qualitative results of the setting without using transformers for feature fusion we use predicted weights for weighted averaging of features via an MLP.
In Tab. 2, we conducted an additional quantitative ablation study w.r.t. the input to the transformer networks. As can be seen, both the projected depth as well as the view ray help the transformer to better fuse the features for the task of 3D reconstruction.
In Fig. 6, we show a comparison of our view selection scheme to the baseline that takes random views from the view candidates (views that contain a specific point). Specifically, we vary the number of views that can be selected. As can be seen, the reconstruction quality gap between the baseline and our method increases with less views which is to be expected since the random selection is more likely to miss important views.
Additional Qualitative Results.
In Fig. 11 further examples are shown that demonstrate the reconstruction capability of our approach. We show a top down, as well as a view from inside the different reconstructed room scenes.
Appendix B Reproducibility
Note that surface reconstruction doesn’t need to be extracted for every frame. It can either be done at the end, when all image features are already fused into the feature volume, or incrementally every couple of frames, on a per-chunk basis, if interactive feedback is desired. In Tab. 4, we report execution times for a chunk of size m. Both, coarse and fine features are spatially refined using a 3D CNN and surface occupancy is computed using the occupancy MLP at a voxel resolution of cm, but only for near-surface voxels, as predicted by coarse and fine near-surface masks. Finally, the mesh is extracted using Marching cubes . Note that our implementation uses high-level PyTorch routines, as well as CPU code (e.g., for Marching Cubes) and, thus, the implementation is not optimized for runtime. A more optimized implementation can be achieved via customized CUDA code. Another interesting avenue towards higher frame rates is the use of sparse 3D convolutions instead of dense 3D convolutions. Feature fusion timings are reported for our default reconstruction setting, when we store views for every feature grid voxel. The feature fusion execution can be further accelerated by using less views. The frames per second (FPS) increase from FPS for views to FPS for and to FPS for views.
Network Architectures.
In Fig. 7, we depict the architectures of the neural networks used in our approach. For both, the coarse and fine layer, we use independent feature fusion and feature refinement networks. The building blocks used in these networks are detailed in Fig. 8.
Reproducibility of Experiments.
To ensure the reproducibility of our experiments, we ran our approach and ablations times. The resulting F-score mean and standard deviation for different experiments is shown is Fig. 9, with standard deviations visualized as error bars.
Appendix C Limitations
In Fig. 10, we present some limitations of our approach. When objects are only partially observed, with a large occluded region, the occluded parts can be missing or lack detail. Examples are missing chair legs, or inaccurate reconstruction of small details such as books. Another challenging scenario for our method is reconstruction of transparent objects, such as glass windows. Most of the time these transparent surfaces get labeled as free-space instead.
Appendix D Data
To train and evaluate our method, we use the ScanNet dataset , which is available under a non-commercial academic licensehttp://kaldir.vc.in.tum.de/scannet/ScanNet_TOS.pdf. ScanNet collects data of static indoor environments, and the ScanNet authors report that consent was obtained from the people whose private spaces were scanned. The ScanNet scenes and locations have been anonymized.