Implicit Functions in Feature Space for 3D Shape Reconstruction and Completion

Julian Chibane, Thiemo Alldieck, Gerard Pons-Moll

Introduction

While many works focus on image-based 3D reconstruction , in this paper, we focus on 3D surface reconstruction and shape completion from a variety of 3D inputs, which are deficient in some respect: low-resolution voxel-grids, high-resolution voxel-grids, sparse and dense point-clouds, complete or incomplete. Such inputs are becoming ubiquitous as 3D scanning technology is increasingly accessible, and they are often an intermediate output of 3D computer vision algorithms. However, the final output for most applications should be a renderable continuous and complete surface, which is the focus of our work.

For sparse grids and (incomplete) point clouds, learning-based methods are a better choice than classical methods , as they reason about global object shape, but are limited by their output representation. Mesh-based methods typically learn to deform an initial convex template , and hence can not represent different topologies. Voxel-based representations have a large memory footprint, which critically limits the output resolution to coarse shapes, without detail. Point cloud representations are more efficient but do not trivially enable rendering and visualization of the surfaces.

Recently, implicit functions have shown to be a promising shape representation for learning. The key idea is to learn a function which, given a coarse shape encoded as a vector, and the x-y-z coordinates of a query point, decide whether the point is inside or outside of the shape. The learned implicit function can be evaluated at query 3D points at arbitrary resolutions, and the mesh/surface can be extracted applying the classical marching cubes algorithm. This output representation enables shape recovery at arbitrary resolutions, is continuous and can handle different topologies.

While these approaches work well to reconstruct aligned rigid objects, we observed they suffer from two main limitations: 1) they can not represent complex objects like articulated humans (reconstructions often miss arms or legs), 2) they do not retain detail present in the input data. We hypothesize this occurs because 1) networks learn an overly strong prior on x-y-z point coordinates damaging the in-variance to articulation, and 2) the shape encoding vector lacks 3D structure, resulting in decodings that look more like classification into shape prototypes rather than continuous regression. Consequently, all existing learning-based approaches, either based on voxels, meshes, points or implicit functions are lacking in some respect.

In this paper we propose Implicit Feature Networks (IF-Nets), which, unlike previous work, do well in 5 different axis, shown in Table 1: they are continuous, can handle multiple topologies, can complete data for sparse input, retaining the nice properties of implicit function models , but crucially they also retain detail when is present in the input (dense input), and can reconstructed articulated humans. IF-Nets differ from recent work in two crucial aspects. First, instead of using a single vector to encode a 3D shape, we extract a 3-dimensional multi-scale tensor of deep features, which is aligned with original Euclidean space embedding the shape. Second, instead of classifying x-y-z point coordinates directly, we classify deep features extracted at continuous query points. Hence, unlike previous work, IF-nets do not memorize common x-y-z locations, which are arbitrary under Euclidean transformations. Instead, they make decisions based on multi-scale features encoding local and global object shape structures around the point.

To demonstrate the advantages of IF-Nets, first, we show that IF-Nets can reconstruct simple rigid 3D objects at better accuracy than previous methods. In ShapeNet , IF-Nets outperform the state-of-the-art results. For articulated humans, we train IF-Nets and related methods on a dataset of 16001600 humans in varied poses, shapes, and clothing. In stark contrast to recent work , IF-Nets can reconstruct articulated objects globally without missing limbs, while recovering detailed structures such as cloth wrinkles. Quantitative and qualitative experiments validate that IF-Nets are more robust to articulations and produce globally consistent shapes without losing fine-scale detail. To encourage further research in 3D processing, learning, and reconstruction, we make IF-Nets publicly available at https://virtualhumans.mpi-inf.mpg.de/ifnets/.

Related Work

Approaches for 3D shape reconstruction can be classified according to the representation used: voxels, meshes, point clouds, and implicit functions; and according to the object type: rigid objects vs humans. For a more exhaustive recent review, we refer the reader to . A condensed overview of strengths and weaknesses of recent 3D reconstruction approaches is given in Table 1.

Voxels for rigid objects: Since voxels are a natural 3D extension to pixels in image grids and admit 3D convolutions, they are most commonly used for generation and reconstruction . However, the memory footprint scales cubically with the resolution, which limited early works to predict shapes in small 32332^{3} grids. Higher resolutions have been used at the cost of limited training batches and slow training or lossy 2D projections . Multi-resolution reconstruction reduced the memory footprint, allowing grids of size 2563256^{3}. However, the approaches are complicated to implement, require multiple passes over the input, and are still limited to grids of size 2563256^{3}, which result in visible quantization artifacts. To smooth out noise, it is possible to represent shapes as Truncated Signed Distance functions for learning . The resolution is however still bounded by the 3D grid storing the TSDF values.

Generative shape models typically map a 1D1D vector to a voxel representation with a neural network . Like us, the authors of observe that the 1D1D vector is too restrictive to generate shapes with global and local structures. They introduce a hierarchical latent code with skip connections. Instead, we propose a much simpler 3-dimensional multi-scale feature tensor, which is aligned with original Euclidean space embedding the shape.

Humans with voxels: From images, CNN based reconstruction of humans represented as voxels or depth-maps typically produce more details than mesh or template-based representations, because predictions are aligned with the input pixels. Unfortunately, this comes at the cost of missing parts in the body. Hence, some methods fit the SMPL model to the reconstructions as a post-processing step. This is however prone to fail if the original reconstructions are too incomplete. All these approaches process image pixels whereas we focus on processing 3D data directly. Unlike our IF-Nets, these methods are bounded by the resolution of the voxel grid.

Meshes for rigid objects: Most mesh-based methods predict shape as a deformation from a template and hence are limited to a single topology. Alternatively, the mesh (vertices and faces) can be inferred directly – while this research direction is promising, methods are still computationally expensive and can not guarantee a closed mesh without intersections. Direct mesh prediction can also be obtained using a learnable version of the classical marching cubes algorithm , but the approach is limited to an underlying small voxel grid of 32332^{3}. Promising combinations of voxels and meshes have been proposed , but results are still coarse.

Meshes for humans: Since the introduction of the (mesh-based) SMPL human model there have been a growing number of papers leveraging it to reconstructing shape and pose from point clouds, depth data and images . Since SMPL does not model clothing and detail, recent methods predict deformations from SMPL or a template . Unfortunately, CNN based mesh predictions tend to be over-smooth. More detail can be obtained predicting normals and displacement maps on a UV-map/geometry image of the surface . However, all these approaches require different templates for every new garment topology or do not produce high-quality reconstructions .

Point clouds for rigid objects: Processing point clouds is an important problem as they are the output of many sensors (LiDAR, 3D scanners) and computer vision algorithms. Due to their low weight, they have been also popular in computer graphics for representing and manipulating shapes . PointNet based architectures were the first to process point clouds directly for classification and semantic segmentation. The idea is to apply a fully connected network to each point followed by a global pooling operation to achieve permutation invariance. Recent architectures apply kernel point convolutions , tree-based graph convolutions , and normalizing flows . Point clouds are also used as shape representation for reconstruction and generation . Unlike voxels or meshes, point clouds need to be non-trivially post-processed using classical methods to obtain renderable surfaces.

Point clouds for humans: Very few works represent humans with point clouds , probably because they can not be rendered. Recent works have employed either PointNet architectures or architectures based on point bases to register a human mesh to the point cloud.

Implicit Functions for rigid objects: Recently, neural networks have been used to learn a continuous implicit function representing shape . For this a neural network can be feed with a latent code and a query point (xx-yy-zz) to predict the TSDF value or the binary occupancy of the point . A recent method achieved state-of-the-art results for 3D reconstruction from images combining 3D query point features with local image features, by approximating the projection of the query point onto the 2D image with a view-point prediction. This trick of querying continuous points used in implicit function learning allows predicting in continuous space (potentially at any resolution), breaking the memory barrier of voxel-based methods. These works inspired our work, but we note that they can not reconstruct articulated humans from 3D data: can not take 3D inputs as point clouds or voxel grids and relies on an approximate 3D to 2D projection losing details; the reconstructions of often miss limbs. We hypothesize that memorize point coordinates instead of reasoning about shape, and that the vectorized latent 1D vector representation is not aligned with the input, and lacks 3D structure. We address this issues by querying deep features extracted at continuous locations of a 3D grid of multi-scale features aligned with the 3D input space. This modification is easy to implement and results in significant gains in reconstruction quality.

Implicit Functions for humans: TSDFs have been used to represent human shapes for depth-fusion and tracking . Such implicit representation has been combined with the SMPL body model to significantly increase tracking robustness and accuracy . From an input image, humans in clothing are predicted using an implicit network . produces higher quality results compared to prior implicit function work . The reconstruction is done by pointwise occupancy prediction based on the location of a 3D query point and 2D image features. For simple poses, the approach produces very compelling and detailed results but struggles for more complex poses. The approach does not incorporate a multi-scale 3D shape representation like ours, and it is designed for image reconstruction, whereas we focus on 3D reconstruction from sparse and dense point clouds and occupancy grids. Like previous implicit networks, our method produces continuous surfaces at arbitrary resolution. But importantly, by virtue of our 3D multi-scale shape representation aligned with the input space, our reconstructions preserve global structure while retaining fine-scale detail, even for complex poses.

Method

To motivate the design of our Implicit Feature Networks (IF-Nets), we first describe the formulation of recent learned implicit functions, pointing out their strengths and weaknesses in Sec. 3.1. We explain our IF-Nets in Sec. 3.2. The key ideas of IF-Nets are illustrated in Fig. 1.

Once f(⋅)f(\cdot) is learned, it can be queried at continuous point locations, without resolution restrictions imposed by typical voxel grids. To construct a mesh, marching cubes can be applied on the predicted occupancy grid. This elegant formulation breaks the barriers of previous representations allowing detailed reconstruction of complex topologies, and has proven effective for several tasks such as rigid object reconstruction from images, occupancy grids, and point clouds. However, we observed that models of this kind suffer from two main limitations: 1) they can not represent complex objects like articulated objects, and 2) they do not preserve detail present in the input data. We address these limitations with IF-Nets.

2 Implicit Feature Networks

We identify two potential problems with the previous formulation. First, directly inputting point coordinates p\mathbf{p} gives the network the option to by-pass reasoning about shape structure, by memorizing typical point occupancies for object prototypes. This severely damages reconstruction in-variance to rotation and translation, which is one of the cornerstones of successful 2D convolution networks for segmentation, recognition, and detection. Second, encoding the full shape in a single vector z\mathbf{z} loses detail present in the data, and loses alignment with the original 3D space where shapes are embedded.

Shape Decoding:

The point encoding F1(p),..,Fn(p)\mathbf{F}_{1}(\mathbf{p}),..,\mathbf{F}_{n}(\mathbf{p}), with Fk(p)∈Fk\mathbf{F}_{k}(\mathbf{p})\in\mathcal{F}_{k}, is then fed into a point-wise decoder f(⋅)f(\cdot), parameterized by a fully connected neural network, to predict if the point p\mathbf{p} lies inside or outside the shape:

In contrast to Eq. (1), in this formulation, the network classifies the point based on local and global shape features, instead of point coordinates, which are arbitrary under rotation, translation, and articulation transformations. Furthermore, due to our multi-scale encoding, details can be preserved while reasoning about global shape is still possible.

3 Method Training

which sums over training surfaces i∈B⊂1,…,Ti\in\mathcal{B}\subset 1,\dots,T of a given mini-batch B\mathcal{B} and point samples j∈R⊂1,…,Sj\in\mathcal{R}\subset 1,\dots,S of a subsample R\mathcal{R}. The subsample R\mathcal{R} is regenerated for every evaluation of the mini-batch loss LB\mathcal{L}_{\mathcal{B}}. For L(⋅,⋅)L(\cdot,\cdot), we use the standard cross-entropy loss. By minimizing LB\mathcal{L}_{\mathcal{B}}, we train the encoder gw(⋅)g_{\mathbf{w}}(\cdot) and the decoder fw(⋅)f_{\mathbf{w}}(\cdot) jointly and end-to-end. Please see the supplementary material for the concrete values for the hyperparameters used in the experiments.

4 Method Inference

Experiments

In this section we validate the effectiveness of IF-Nets on the challenging task of 3D shape reconstruction. We show that our IF-Nets are able to address two limitations of recent learning-based approaches for this task: 1) IF-Nets preserve detail present in the input data, while also reasoning about incomplete data, 2) IF-Nets are able to reconstruct articulated humans in complex clothing. To this end, we conduct three experiments of increasing complexity: Point Cloud Completion (Sec. 4.1), Voxel Super-Resolution (Sec. 4.2) and Single-View Human Reconstruction (Sec. 4.3).

Baselines: For the task of Point Cloud Completion, we evaluate our approach against Occupancy Networks (OccNet), Point Set Generation Networks (PSGN) and Deep Marching Cubes (DMC). For Voxel Super-Resolution, we compare against IMNET as well as again against OccNet and DMC. For DMC and PSGN we used the implementations provided online by the authors of . We trained all methods until the validation minimum was reached. Training was repeated for every considered experiment setup. To show a consistent comparison, we modified the IMNET implementation to be able to be trained on all ShapeNet classes jointly. For IMNET and OccNet, we kept the sampling strategies proposed by their authors. For IMNET, we followed the authors and performed progressive resolution increasing of training data sampling during training.

Metrics: To measure reconstruction quality quantitatively, we consider three established metrics (see suppl. of for definition and implementation): volumetric intersection over union (IoU) measuring how well the defined volumes match (higher is better), Chamfer-L2L_{2} measuring the accuracy and completeness of the surface (lower is better), and normal consistency measuring the accuracy and completeness of the shape normals (higher is better).

Data: We consider two datasets: 1) a dataset containing 3D scans of humansThe dataset will be available for purchase from Twindom. to assess the challenging task of reconstruction from incomplete and articulated shapes and 2) the established ShapeNet dataset, consisting of rigid object classes, with rather prototypical shapes like cars, airplanes, and rifles. The ShapeNet data has been pre-processed to be watertight by the authors of , allowing to compute ground truth occupancies and scaled such that every shape’s largest bounding box edge has length one. We conduct all experiments and evaluations using pre-possessed ShapeNet data and use the common training and test split by . However, preprocessing failed for some objects, leading to broken objects with large holes. Therefore, 508 heavily distorted objects have been removed for meaningful evaluation. The filtered list of all used objects is published alongside the code. We also evaluate on a challenging dataset consisting of scanned humans in highly varying articulations with complex and varying clothing topologies like coats, skirts, or hats. The scans have been captured using commercial 3D scanners. The dataset, referred to as Humans, consists of 2183 such scans, split into 478 examples for testing, 1598 for training and 197 for validation. The scans have been height normalized and centered, but in contrast to the ShapeNet objects, exhibit varying rotations.

As a first task, we apply IF-Nets to the problem of completing sparse and dense point clouds – we sample 300 points (sparse) and 3000 points (dense) respectively from ShapeNet surface models and ask our method to complete the full surfaces. Completing point clouds is challenging since it requires to simultaneously preserve input details and reason about missing structure at the same time. In Fig. 2 we show comparisons against the baseline methods. Our method outperforms all baselines both in preserving local detail and recovering global structures. For the dense point clouds, the strengths of our method are paramount. Our method is the only one capable of reconstructing the car rear-view mirrors and the additional shelf of the wardrobe. We additionally quantitatively compare our method and report the numbers in Tab. 2. Our method beats the state-of-the-art in all metrics by large margin. In fact, using 3000 points as input, all competitors produce results which have larger Chamfer distance than the input itself, suggesting they fail at preserving input detail. Only IF-Nets preserve input details while completing missing structures.

2 Voxel Super-Resolution

As a second task, we apply our method to 3D super-resolution. To effectively solve this task, our method needs to again preserve the input shape while reconstructing details not present in the input. Our results in side-by-side comparison with the baselines are depicted in Fig. 2 (bottom). While most baseline methods either hallucinate structure or completely fail, our method consistently produces accurate and highly detailed results. This is also reflected in the numerical comparison in Tab. 3, where we improve over the baselines in all metrics.

The two last examples in Fig. 2 illustrate the limitations of current implicit methods: If a shape differs too much from the training set, the method fails or seems to return a similar previously seen example. Consequentially, we hypothesize that the current methods are not suited for tasks where classification into shape prototypes is not sufficient. This is for example the case for humans as they come in various shapes and articulations. To verify our hypothesis, we additionally perform 3D super-resolution on our Humans dataset. Here the advantages are even more prominent: Our method is the only one that consistently reconstructs all limbs and produces highly detailed results. Implicit learning-based baselines produce truncated or completely missing limbs. We outperform all baselines also quantitatively (see Tab 4).

3 Single-View Human Reconstruction

Finally, to demonstrate the full capabilities of IF-Nets, we use them for single-view human reconstruction. In this task, only a partial 3D point cloud is given as input – the typical output of a depth camera. We conduct this experiment on the challenging Humans dataset, by rendering a 250×250250\times 250 resolution depth image, yielding around 5000 points on the visible side of the subject. To successfully fulfill this task, our model has to simultaneously reconstruct novel articulations, retain fine details, and complete the missing data at the occluded regions – the input contains only one side of the underlying shape. Despite these challenges, our model is capable of reconstructing plausible and highly detailed shapes. In Fig. 4, we show both input and our results from four different angles. Note how fine structures like the scarfs, wrinkles, or individual fingers are present in the reconstructed shapes. Although the backside region (occluded) has less details than the visible one, IF-Nets always produce plausible surfaces.

This can also be seen quantitatively. Ours: IoU 0.86, Chamfer-L2L_{2} 0.011×10−20.011\times 10^{-2}, Normal-Consistency 0.90. Input point cloud: Chamfer-L2L_{2} 0.252×10−20.252\times 10^{-2}. The quantitative results are in between the reconstruction quality of 32332^{3} and 1283128^{3} full subject voxel inputs (see Tab. 4), which once more validates that IF-Nets can complete single-view data. In the supplementary video, we show an additional result on single-view reconstruction on the BUFF dataset from video (without retraining nor fine tuning the model).

Discussion and Conclusion

In this work, we have introduced IF-Nets for 3D reconstruction and completion from deficient 3D inputs. First, we have argued for an encoding consisting of a 3D multi-scale tensor of deep features, which is aligned with the Euclidean space embedding the shape. Second, instead of classifying x-y-z coordinates directly, we classify deep features extracted at their location. Experiments demonstrate that IF-Nets deliver continuous outputs, can reconstruct multiple topologies such as 3D humans in varied clothing, and 3D objects from ShapeNet. Quantitatively, IF-Nets outperform all state-of-the-art baselines by a large margin in all tasks. Our reconstruction from single-view point clouds(detailed on the visible part but with missing data on the occluded part), demonstrate the strengths of IF-Nets: details in the input are preserved, while the shape is completed on the occluded part, even for articulated shapes.

Future work will explore extending IF-Nets to be generative, that is being able to sample detailed hypothesis conditioned on partial input. We also plan to address image-based reconstruction in 2 stages: first predicting a depth map, and then completing shape with IF-Nets.

With a rising number of computer vision image reconstruction methods producing partial 3D point clouds and voxels, and 3D scanners and depth cameras becoming accessible, 3D (deficient and incomplete) data will be omnipresent in the future, and IF-Nets have the potential to be an important building block for its reconstruction and completion.

Acknowledgments. We would like to thank Verica Lazova for helping creating the figures, Bharat Lal Bhatnagar for helping with data preprocessing, and Lars Mescheder for sharing their watertight Shapenet meshes. This is work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans). We would like to thank Twindom for providing us with the scan data.

References