Playing for 3D Human Recovery
Zhongang Cai, Mingyuan Zhang, Jiawei Ren, Chen Wei, Daxuan Ren, Zhengyu Lin, Haiyu Zhao, Lei Yang, Chen Change Loy, Ziwei Liu
Introduction
Image- and video-based 3D human recovery, i.e., simultaneous estimation of human pose and shape via parametric models such as SMPL , have transformed the landscape of holistic human understanding. This technology is critical for entertainment, gaming, augmented and virtual reality industries. However, despite that the exciting surge of deep learning is arguably driven by enormous labeled data , the same is difficult to achieve in this field. The insufficiency of data (especially in the wild) is attributed to the prohibitive cost of 3D ground truth (particularly parametric model annotation) . Existing datasets are either small in scale , collected in constrained indoor environment , or not providing the 3D parametric model annotation at all .
Inspired by the success of training deep learning models with video game-generated data for various computer vision tasks such as instance segmentation , 2D keypoint estimation , motion prediction , mesh reconstruction , detection and tracking , we present GTA-Human in the hope to address the aforementioned limitations of existing datasets. GTA-Human is built by coordinating a group of computational workers (Figure 2) that simultaneously play the popular video game Grand Theft Auto V (GTA-V), to put together a large-scale dataset (Table I) with 1.4 million SMPL parametric labels automatically annotated in 20 thousand video sequences. Besides the scale, GTA-Human explores the rich resources of the in-game database to diversify the data distribution that is challenging to achieve in real life (Figure 3, 4 and 5): more than 600 subjects of different gender, age, ethnicity, body shape and clothing; 20,000 action clips comprising a wide variety of daily human activities; six major categories of locations with drastically different backgrounds from city streets to the wild; camera angles are manipulated in each sequence to reflect a realistic distribution; subject-environment interaction that gives rises to occlusion of various extents; time of the day that affects lighting conditions, and weather system that mimics the real climate changes.
Equipped with GTA-Human, we conduct an extensive investigation in the use of synthetic data for 3D human recovery. 1) Better 3D human recovery with data mixture. Despite the seemingly unavoidable domain gaps, we show that practical settings that mix synthetic data with real data, such as blended training and pretraining followed by finetuning, are surprisingly effective. First, HMR , one of the first deep learning-based methods for SMPL estimation with relatively simplistic architecture, when trained with data mixture, is able to outperform more recent methods with sophisticated designs or additional information such as SPIN and VIBE . Moreover, PARE , a state-of-the-art method also benefit considerably from GTA-Human. Second, our experiments on the video-based method VIBE further demonstrate the effectiveness of data mixture: an equal amount of synthetic GTA-Human data is as good as a real-captured indoor dataset as the frame feature extractor is already pretrained on real datasets; the full set of GTA-Human is even on par with in-domain training data.
2) Closing the domain gap with synthetic data. We then study the reasons behind the effectiveness of game-playing data. An investigation into the domain gaps provides insights into the complementary nature of synthetic and real data: despite the reality gap, the synthetic data embodies the diversity that most of the real data lack, as the latter is typically collected indoors. Moreover, we experiment with mainstream domain adaptation methods to further close the domain gaps and obtain improvements.
3) Dataset scale matters. We demonstrate that adding game-playing data progressively improves the model performance. Considering the difficulty of collecting real data with ground truth 3D annotations, synthetic data may thus be an attractive alternative. Moreover, a multi-factor analysis reveals that supervised learning leads to severe sensitivity to data density. Amongst factors such as camera angles, pose distributions, and occlusions, a consistent drop in performance is observed where data is scarce. Hence, our observation suggests synthetic datasets may play a vital role to supplement corner case scenarios in the future.
4) Strong supervision (SMPL) is key. Compared to large-scale pose estimation benchmarks that only provide 3D keypoints, we demonstrate that strong supervision in the form of SMPL parameters may be quintessential for training a strong model. We discuss the potential reasons behind this observation, which reaffirms the value of GTA-Human as a scalable training source with SMPL annotations.
5) Big data benefits big models. Despite recent development in deeper convolutional networks and vision transformers in computer vision research, the mainstream backbone size remains unchanged for 3D human recovery . We extend our study to deeper CNNs and Transformers, and show that training with GTA-Human not only gives rise to improvements but allows smaller backbones to outperform larger counterparts.
Related Work
Registration-based Methods. As the output of the human parametric model is manipulated by body parameters, SMPLify and the following SMPLify-X are the pioneering works to optimize these parameters to minimize the distance between ground truth 2D keypoints and reprojected human mesh joints. SMPLify is also extended to videos with temporal constraints employed . Although optimization-based methods are able to achieve impressive results, they are slow and typically take more than 60 seconds per frame. Hence, recent work has been proposed to accelerate optimization.
Regression-based Methods. Direct regression of body parameters using a trained deep learning model has gained more popularity due to fast inference. The recent works are categorized into image-based , and video-based methods. HMR is a pioneering end-to-end deep learning-based work, which takes ResNet-50 as its backbone and directly regress the parameters of and . VIBE is a milestone video-based work that leverages temporal information for realistic pose sequences. Recently, a transformer encoder is introduced for vertex-joint reweighting , but the method still uses a CNN backbone for feature extraction.
Mixed Methods. There is a line of work that combines optimization-based and regression-based techniques. SPIN , adds a SMPLify step to produce pseudo parametric labels to guide the learning of the network. SPIN address the lack of SMPL annotation but the optimization step results in slow training. Others propose to refine the per-frame regression results by bundle adjustment of the video sequence as a whole , designing a new swing-twist representation to replace the original axis-angle representation of SMPL , and finetuning a trained network to obtain refined prediction , and employing a network to predict a parameter update rule in iterations of optimization .
2 Datasets
Datasets with 2D Keypoint Annotations. Many datasets contain in-the-wild images, albeit the lack of SMPL annotations, they provide 2D keypoint labels. Datasets such as LSP , LSP-Extended , COCO and MPII contain images crawled from the Internet, and are annotated with 2D keypoints manually. Such a strategy allows a large number of in-the-wild images to be included in the dataset. To obtain 3D annotations that are crucial to human pose and shape estimation, a common method is to fit an SMPL model on 2D keypoints. SSP-3D and 3DOH50K leverages pre-trained model to perform keypoint estimation as the first step, whereas UP-3D and EFT performs fitting on ground truth keypoints. However, these datasets typically suffer from the inherent depth ambiguity of images and the pseudo-SMPL may not have the accurate scale.
Real Datasets. Motion capture facilities are built to achieve high-accuracy 3D annotations. HumanEva and Human3.6M employ optical motion capture systems, but intrusive markers are needed to be placed on the subjects. Total Capture MuPoTS-3D , Panoptic Studio , and HUMBI make use of multiple camera views and require no intrusive marker. However, the background is constant and thus lacks diversity. 3DPW combines inertial measurement units (IMUs) and a moving camera to build an in-the-wild dataset with 3D annotations. 3DPW has become an important benchmark for 3D human recovery. Nevertheless, the IMU drift is still an obstacle and the dataset only contains a relatively small number of videos. SMPLy constructs point clouds from multi-view capture of static people and fits SMPL on them. However, the scale of the dataset is limited by the difficulty of collecting videos that meet the special setup requirement. HuMMan is the most recent large-scale multi-modal 4D human dataset.
Synthetic or Mixed Datasets. SURREAL , Hoffmann et al. render textured SMPL body models in real-image backgrounds. However, this strategy does not account for the geometry of the clothes, where the mismatch may result in unrealistic subjects. 3DPeople uses clothed human models while MPI-INF-3DHP takes segmented subjects from images and paste them onto new backgrounds in the training set. However, the subject-background interaction is still unnatural. AGORA is a recent synthetic dataset featuring high-quality annotations by rendering real human scans in a virtual world. However, the dataset is image-based and does not support the training of video-based methods. Richter et al. , Krähenbühl et al. , JTA , GTA-IM , SAIL-VOS 3D , MOTSynth have demonstrated the potential of obtaining nearly free and perfectly accurate annotations from video games for various computer vision tasks. However, these datasets do not provide SMPL annotation needed for our investigation. We take inspirations from these works in building GTA-Human.
GTA-Human Dataset
A scale comparison between GTA-Human and existing dataset is shown in Table I. GTA-Human features 1.4 million individual SMPL annotations, which is highly competitive compared to other real datasets and synthetic datasets with realist setups. Moreover, GTA-Human consists of in-the-wild scenes that are expensive and difficult to collect in real life. Notably, GTA-Human provides video sequences instead of static frames and supports video-based human recovery.
Inspired by existing works that use GTA-generated data for various vision tasks , our toolchain extracts ground truth 2D and 3D keypoints, semantic and depth maps from the game engine, followed by fitting SMPL models on the keypoints with temporal constraints. To achieve scalability and efficiency, we design and deploy an automatic system (Fig. 2) that leverages cloud-based services for parallel deployment and coordination of our tools on a large number of computer instances and GPU cluster nodes. GTA-Human consists of sequences of single-person scenes. More examples in GTA-Human are found in Fig. 6.
Cloud-based NoSQL Database. The unit of data in GTA-Human is a single video sequence. Hence, we employ a Database that is hosted on the cloud, to track the progress of data generation and processing of each sequence. The status of a sequence is updated at each stage in the toolchain, which we elaborate on the details below.
Scenario File Generator. This tool reads from the Database to retrieve sequence IDs that are either not generated before or failed in previous processing attempts, and produce random scene attributes such as subject ID, action ID, location in the 3D virtual world, camera position and orientation, lighting, and weather settings.
Cloud-based Message FIFO Queue. The Message FIFO Queue parse the scenario files from the Scenario File Generator as text strings, which can be fetched in the first-in-first-out (FIFO) manner by multiple Local GUI Workers. Note that the queue allows for multiple workers to retrieve their next jobs simultaneously.
Local GUI Workers. We purchase multiple copies of GTA-V and install them on regular gaming desktops. We refer to these desktops as Local GUI Workers. Each worker runs three tools: Scenario Controller, Data Collector, and Data Analyser that we elaborate on below.
Scenario Controller. Taking scenario files as the input, Scenario Controller is essentially a plugin that interacts with the game engine via the designated Application Programming Interface (API). It is thus able to control the subject generation and placement, action assignment to the subject, camera placement, in-game time, and weather.
Data Collector. This tool obtains data and some annotations from the API provided by GTA-V. First, it extracts 3D keypoints from each subject via the API provided by GTA-V. In addition to the original 98 keypoints available, we further obtain head top and nose from interpolation of existing keypoints. We project 3D keypoints to the image plane with known intrinsic and extrinsic parameters of the camera to obtain 2D keypoints. Second, we project light rays at each joint to determine if the joint is occluded or self-occluded by checking the entity that the light ray hits first . Third, our tool intercepts the rendering pipeline, powered by DirectX, for depth maps and semantic masks. The pixel-wise depth is directly read from depth buffers. Shader injection enables the segmentation of individual patches, and we manually assign the semantic class to various shaders based on their variable names. We refer interested readers to for more details. Fourth, the collector also records videos.
Data Analyser. To filter out low-quality data in the early stage, Data Analyser imposes several constraints on 3D keypoints obtained. We compute joint movement speed simply as the position different in consecutive frames to filter out less expressive actions (slow-moving or stationary actions). Severely occluded, or out-of-view subjects are also flagged at this stage. If sequences pass the analysis, their data are transferred from the local storage to a centralized storage space on our GPU cluster (Cluster Storage) for further processing. The failed ones, however, are deleted. The Database is notified of the result to get the status updated.
Cluster Workers and SMPL Annotator. On each Cluster Worker (a GPU in the cluster), we run an instance of SMPL Annotator that takes keypoint annotation from the Cluster Storage. We upgrade SMPLify in two ways to obtain accurate SMPL annotation. 1) we find out that compared to 2D keypoints that have inherent depth ambiguity, exacerbated by weak perspective projection , 3D keypoints are unambiguous. Minor modifications are needed to replace the 2D keypoint loss of the original SMPLify with 3D keypoint loss. 2) Taking advantage of the fact that GTA-Human consists of video sequences instead of unrelated images, temporal consistency in the form of rotation smoothing and unified shape parameters are enforced. The SMPL parameters include and , and an additional translation vector, are optimized at an average of one second per frame. We visualize more examples in GTA-Human that are produced with our SMPL annotation tool in Fig. 6.
2 Data Diversity
Due to the difficulty and cost of data collection and SMPL annotation for the 3D human recovery task, most existing datasets are built at restrictive locations such as indoor studios or laboratory environments. Furthermore, only a small number of subjects are usually employed to perform a limited set of actions. In contrast, GTA-Human is designed to maximize the variety in the following aspects. We demonstrate the diversity in subjects, locations, weather, and time (light conditions) in Fig. 3, actions in Fig. 4, and camera angles in Fig. 5.
Subjects. GTA-Human collects over 600 subjects of different genders, ages, ethnicities, clothing, and body shapes for a wide coverage of human appearances. In addition, unlike motion capture systems in real life that rely on intrusive markers to be placed on the subjects, accurate skeletal keypoints are obtained directly from the game’s API.
Actions. Existing datasets either design a small number of actions , or lack a clearly defined action set . In contrast, we gain access to a large database of motion clips (actions) that can be used to manipulate the virtual characters, whose typical length is 30-80 frames at 30 FPS. These actions provide a fairly holistic representation of city-dwellers’ daily activities, and are reasonably realistic because they are originally produced via motion capture of real human actors or actresses. We select 20,000 most dynamic and expressive actions. In Fig. 4, the distribution of GTA-Human poses does not only have the widest spread, but also covers existing poses in the real datasets to a large extent. Note that these actions allows for the study on video-based methods in Section 4.2.
Locations. The conventional optical or multi-view motion capture systems require indoor environments, resulting in the scarcity of in-the-wild backgrounds. Thanks to the open-world design of GTA, we have seamless access to various locations with diverse backgrounds, from city streets to the wilderness. Our investigation in Section 4.2 highlights these diverse locations are complementary to real datasets that are typically collected indoor.
Camera Angle. Recent studies have shown the critical impact of camera angles on model performance, yet its effect in 3D human recovery is not fully explored due to data scarcity: it is common to have datasets with fixed camera positions . In GTA-Human, we choose to sample random camera positions from the distribution of the real datasets to balance both diversity and realness. Our data collection tool enables the control of camera placement position and orientation, thus allowing the study on camera angles that is otherwise difficult in real life. We visualize the camera angles in Fig. 5.
Interaction. Compared to existing works that crop and paste subjects onto random backgrounds , the subjects in GTA-Human are rendered together with the scenes to achieve a more realistic subject-environment interaction empowered by the physics engine. Interesting examples include the subject falling off the edge of a high platform, and the subject stepping into a muddy pond causing water splashing. Moreover, taking advantage of the occlusion culling mechanism , we are able to annotate the body joints as “visible” to the camera, “occluded” by other objects, or “self-occluded” by the subject’s own body parts.
Lighting and Weather. Instead of adjusting image exposure to mimic different lighting, we directly control the in-game time to sample data around the clock. Consequently, GTA-Human contains drastically different lighting conditions and shadow projections. We also introduce random weather conditions such as rain and snow to the scenes that would be otherwise difficult to capture in real life.
Experiments
In this section, we study how to use game-playing data for 3D human recovery for real-life applications.
Datasets. We follow the original training convention of our baseline methods . we define the “Real” datasets used in the experiments to include Human3.6M (with SMPL annotations via MoSh ), MPI-INF-3DHP , LSP , LSP-Extended , MPII and COCO . “Real” datasets consist of approximately 300K frames. ”Blended” datasets are formed by simply mixing GTA-Human data with the “Real” data. Amongst the standard benchmarks, 3DPW has 60 sequences (51k frames) of unconstrained scenes. In contrast, MPI-INF-3DHP has only two sequences of real outdoor scenes (728 frames) and Human3.6M is fully indoor. Hence, we follow the convention to evaluate models mainly on 3DPW test set to gauge their in-the-wild performances. Nevertheless, we also provide experiment results on Human3.6M and MPI-INF-3DHP.
Metrics. The standard metrics are Mean Per Joint Position Error (MPJPE), and Procrustes-aligned Mean Per Joint Position Error (PA-MPJPE), i.e., MPJPE evaluated after rigid alignment of the predicted and the ground truth joint keypoints, both in millimeters (). We highlight that PA-MPJPE is the primary metric , on which we conduct most of our discussions.
Training Details. We follow the original paper in implementing baselines on the PyTorch-based framework MMHuman3D . HMR+ is a stronger variant of the original HMR, for which we remove all adversarial modules from the original HMR for fast training and add pseudo SMPL initialization (“static fits”) for keypoint-only datasets following SPIN without further in-the-loop optimization. For the Blended Training (BT), since GTA-Human has a much larger scale than existing datasets, we run all our experiments on 32 V100 GPUs, with the batch size of 2048 (four times as SPIN ). The learning rate is also scaled linearly by four times to 0.0002. The rest of the hyperparameters are the same as SPIN . For the Finetuning (FT) experiments, we use the learning rate of 0.00001 with the batch size of 512, on 8 V100 GPUs for two epochs.
Domain Adaptation Training Details. We use the same training settings as Blended Training, except that an additional domain adaptation loss is added in training. CycleGAN , we first train a CycleGAN between real data and our synthetic GTA-Human data. Then we use a trained sim2real generator from the CycleGAN to transform the input GTA-Human image into a real-style image during training. For JAN , we use the default Gaussian kernel with a bandwidth 0.92, and set its loss weight to 0.001. For Chen et al. , we use the default trade-off coefficient 0.1, and set its loss weight to 1e-4. For Ganin et al. , we use a 3-layer MLP to classify the domain of given features extracted from the backbone. The loss weight of the adversarial part is progressively increased to 0.1 for more stable training.
2 Better 3D Human Recovery with Data Mixture
Despite that GTA-Human features reasonably realistic data, there inevitably exists domain gaps. Surprisingly, intuitive methods of data mixture are effective despite the domain gaps for both image- and video-based 3D human recovery.
Image-based 3D Human Recovery. We evaluate the use of synthetic data under two data mixture settings: blended training (BT) and finetuning (FT). Results are collated in Table II. In blended training (BT), synthetic GTA-Human data is directly mixed with a standard basket of real datasets (Human3.6M , MPI-INF-3DHP , LSP , LSP-Extended , MPII and COCO ). Compared with the HMR and HMR+ baselines, blended training achieves 7.0 mm and 5.7 mm improvements in PA-MPJPE, surpassing methods such as SPIN that requires online registration or VIBE that leverages temporal information. As for finetuning (FT), we finetune a pretrained model with mixed data. Since finetuing is much faster than blended training, this allows us to perform data mixture on more base methods such as SPIN and PARE . Finetuning leads to considerable improvements in PA-MPJPE compared to the original HMR (11.8 mm), HMR+ (6.2 mm), SPIN (7.2 mm) and PARE (4.1 mm) baselines.
Video-based 3D Human Recovery. In Table III, we validate that data mixture is also effective for video-based methods. We conduct the study with the popular VIBE as the base model. VIBE uses a pretrained SPIN model as the feature extractor for each frame, and we train the temporal modules with datasets indicated in Table III. We obtain the following observations. First, when training alone, GTA-Human outperforms MPI-INF-3DHP with an equal number of training data. Second, the full set of GTA-Human is comparable with the in-domain training source (3DPW train set), even slightly better in PA-MPJPE. Third, GTA-Human is complementary to real datasets as blended training leads to highly competitive results in all metrics.
Comparison with Other Data-driven Methods. We highlight that GTA-Human is a large-scale, diverse dataset for 3D human recovery. In Table IV, we compare GTA-Human with several other recent works that provide additional data for human pose and shape estimation. We show that GTA-Human is a practical training source that improves the performance of various base methods. Notably, GTA-Human slightly surpasses AGORA, which is built with expensive industry-level human scans of high-quality geometry and texture. This result suggests that scaling with game-playing data at a lower cost achieves a similar effect.
3 Closing the Domain Gap with Synthetic Data
After obtaining good results under both image- and video-based settings on 3DPW, an in-the-wild dataset and the standard test benchmark, we extend our study to answer why is game playing data effective at all? To this end, we also evaluate models on other (mostly) indoor benchmarks such as Human3.6M (Protocol 2) and MPI-INF-3DHP in Table V. Interestingly, we notice that the performance gains on these two benchmarks are not as significant as those on 3DPW. Moreover, existing methods commonly include in-the-wild COCO data in the training set, in addition to popular training datasets are typically collected indoors. We aim to explain the above-mentioned observations and practices through both qualitative and quantitative evaluations.
In Fig. 7, we visualize the feature distribution of various datasets. We discover that there are indeed some domain gaps between real indoor data and real outdoor data. Hence, models trained on real indoor data may not perform well in the wild. We observe that in Fig. 7(a), indoor data has a significant domain shift away from in-the-wild data. This result implies that models trained on indoor datasets may not transfer well to in-the-wild scenes. In Fig.7(b), blended training achieves better results as 3DPW test data are well-covered by mixing real data or GTA-Human data. Specifically, even though the domain gap between GTA-Human and real datasets persists, the distribution of 3DPW data is split into two main clusters, covered by GTA-Human and real datasets separately. Hence, this observation may explain the effectiveness of GTA-Human: albeit synthetic, a large amount of in-the-wild data provides meaningful knowledge that is complementary to the real datasets.
Moreover, we further validate the synergy between real and synthetic data through domain adaptation in Table VI. We select and implement several mainstream domain adaption methods , and evaluate them on an HMR model under BT with an equal amount of real data and GTA-Human (). We discover that learned data augmentation such as CycleGAN may not be effective, whereas domain generalization techniques (JAN and Ganin et al. ) and domain adaptive regression such as Chen et al. further improves the performance. In Fig.7(c), domain adaptation (Ganin et al. ) pulls the distributions of both real and GTA-Human data together, and they jointly establish a better-learned distribution to match that of the in-the-wild 3DPW data.
4 Dataset Scale Matters
We study the data scale in two aspects. 1) Different amounts of GTA-Human data are progressively added in the training to observe the trend in the model performance. 2) The influence of a lack of data from the perspectives of critical factors such as camera angle, pose, and occlusion.
Amount of GTA-Human Data. In Fig. 8, we delve deeper into the impact of data quantity on 3DPW test set. HMR is used as the base model with BT setting. The amount of GTA-Human data used is expressed as multiples of the total quantity of real datasets (300K ). For example, means the amount of GTA-Human is twice as much as the real data in the BT. A consistent downward trend in the errors (6 mm decrease) with increasing GTA-Human data used in the training is observed. Since real data is expensive to acquire, synthetic data may play an important role in scaling up 3D human recovery in real life.
Synthetic Data as a Scalable Supplement. We collate more experiments with different real-synthetic data ratios in Table VII, using HMR+ as the base method and BT as the data mixture strategy. We observe that 1) Adding more data, synthetic and real alike, generally improve the performance. 2) Mixing 75% real data with 25% synthetic data performs well (200K to 400K data). 3) When the data amount increases, high ratio of real data cannot be sustained beyond 300K data due to insufficient real data. However, additional synthetic data still improves model performance. These experiments reaffirm that synthetic data complements real data, and more importantly, synthetic data serves as an easily scalable training source to supplement typically limited real data.
Impact of Data Scarcity. In Fig. 9, we systematically study the HMR+ model trained with BT and evaluate its performance on GTA-Human, subjected to different data density for factors such as camera angle, pose, and occlusion. We discretize all examples evaluated to obtain and plot the data density with bins, and compute the mean error for each bin to form the curves. A consistent observation across factors is that the model performance deteriorates drastically when data density declines, indicating high model sensitivity to data scarcity. Hence, strategically collected synthetic data may effectively supplement the real counterpart, which is often difficult to obtain.
5 Strong Supervision is Key
Due to the prohibitive cost of collecting a large amount of SMPL annotations with a real setup, it is appealing to generate synthetic data that is automatically labelled. In this section, we investigate the importance of strong supervision and discuss the reasons. We compare weak supervision signals (i.e., 2D and 3D keypoints) to strong counterparts (i.e., SMPL parameters), and find out that the latter is critical to training a high-performing model. We experiment under BT setting on 3DPW test set, Table VIII shows that strong supervision of SMPL parameters and , are much more effective than weak supervision of body keypoints. Our findings are in line with SPIN . SPIN tests fitting 3D SMPL on 2D keypoints to produce pseudo SMPL annotations during training and finds this strategy effective. However, this conclusion still leaves the root cause of the effectiveness of 3D SMPL unanswered, as recent work suggests that 2D supervision is inherently ambiguous . In this work, we extend the prior study on the supervision types by adding in 3D keypoints as a better part of the weak supervision and find out that SMPL annotation is still far more effective.
As for reasons that make strong supervision (SMPL parameters) more effective than weak supervision (keypoints), we argue that keypoints only provide partial guidance to body shape estimation (bone length only), but is required in joint regression from the parametric model. Moreover, ground truth SMPL parameters is directly used in the loss computation with the predicted SMPL parameters (Equation 1), which initiates gradient flow that reaches the learnable SMPL parameters in the shortest possible route. On the contrary, the 3D keypoints are obtained with joint regression of canonical keypoints with estimated body shape , and the global rigid transformation derived from the SMPL kinematic tree (Equation 2). The 2D keypoints further require extra estimation of translation for the transformation of the 3D keypoints, and 3D to 2D projection with assumed focal length as well as camera center . The elongated route and uncertainties introduced in the process to compute the loss for 2D keypoints (Equation 3) hinder the effective learning.
6 Big Data Benefits Big Models
ResNet-50 remains a common backbone choice, since HMR is firstly introduced for deep learning-based 3D human recovery. In this section, we extend our study of the impact of big data on more backbone options, including deeper CNNs such as ResNet-101 and 152 , as well as DeiT , as a representative of Vision Transformers. In Table IX, we evaluate various backbones for the HMR baseline. We highlight that including GTA-Human always improves model performance by a considerable margin, regardless of the model size or architecture. Note that using Transformers as the feature extractor for human pose and shape estimation is under-explored in recent literature; there may be some room for further improvement upon our attempts presented here. Nevertheless, the same trend holds for the two transformer variants. Interestingly, additional GTA-Human unleashes the full power of a small model (e.g., ResNet-50), enabling it to outperform a larger model (e.g., ResNet-152) trained with real data only. This suggests data still remains a critical bottleneck for accurate human pose and shape estimation.
Conclusion
In this work, we evaluate the effectiveness of synthetic game-playing data in enhancing human pose and shape estimation especially in the wild. To this end, we present GTA-Human, a large-scale, diverse dataset for 3D human recovery. Our experiments on GTA-Human provide five takeaways: 1) Training with diverse synthetic data (especially with outdoor scenes) achieves a significant performance boost. 2) The effectiveness is attributed to the complementary relation between real and synthetic data. 3) The more data, the better because model performance is highly sensitive to data density. 4) Strong supervision such as SMPL parameters are essential to training a high-performance model. 5) Deeper and more powerful backbones also benefit from a large amount of data. As for future works, we plan to investigate beyond 1.4M data samples with more computation budgets to explore the boundary of training with synthetic data. Moreover, it would be interesting to study the sim2real problem for 3D parametric human recovery more in-depth with GTA-Human, or even extend the game-playing data to other human-related topics such as model-free reconstruction that are out of the scope of this work.
Acknowledgments
This work is supported by NTU NAP, MOE AcRF Tier 2 (T2EP20221-0033), and under the RIE2020 Industry Alignment Fund - Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).
As we are aware that GTA is not a perfect depiction of real life, we address some ethical concerns and explain our strategies to alleviate potential negative impact.
The subjects present in GTA-Human are virtual humans extracted from the in-game database, that usually do not have clear real-life references. Protagonists may have some sort of real-life references, but the appearances are altered to suit the corresponding characters in the context of the game storyline.
Violence and sexualized actions
We manually screen around 1k actions from the 20k classes and find that the vast majority of the actions are used to depict the ordinary lives of the city-dwellers (e.g. walking, drinking coffee, doing push-ups, and so on). This is further supported by 1) the distribution of GTA-Human is center-aligned with real datasets. 2) methods trained on GTA-Human can perform convincingly better on standard real datasets. Both indicate that the domain shift in actions may not be noticeably affected by the small portion of offensive actions. Moreover, no weapon is depicted in the dataset.
Stereotypes and Biases
The original storylines of GTA-V may insert strong stereotypes in the depiction of characters depending on their attributes (e.g., gender and race). We thus adopt the following strategies to minimize biases.
First, all factors including the characters and actions, are decoupled and randomized in GTA-Human. Specifically, the examples in GTA-Human are not linked to the original storylines; all the characters and the actions are pulled out of the in-game database, randomly assigned at random locations all over the map. Hence, it is very unlikely any character-specific actions could be reproduced.
Second, we have conducted a manual analysis on clothing (which may serve as an indication of social status) vs gender and race. As we find it difficult to determine if a specific attire has certain social implications without context (for example, skin-showing attire not be associated with sex workers as it is also common to find people in bikinis at the beach), we thus categorize all clothing into formal, semi-formal and casual. We observe that while there is approximately the same number of men and women in formal attire (11% vs 9%); more men in casual attires (e.g. tank tops, topless) than women (25% vs 11%); all races have approximately the same distribution of formal, semi-formal and casual (). Hence, we find the character appearances mostly (albeit not perfectly) balanced across genders and races.
Third, as much as we hope to perform a complete and thorough data screening and cleaning, we highlight it is not very practical to manually inspect all examples due to the sheer scale of the dataset. Hence, we anonymize the characters, actions, and locations such that they exist in the dataset to enrich the distribution, but cannot be retrieved for malicious uses.
Copyright
The publisher of GTA-V allows for the use of game-generated materials provided that it is non-commercial, and no spoilers are distributedPolicy on mods. http://tinyurl.com/yc8kq7vn.Policy on copyrighted material. http://tinyurl.com/pjfoqo5.. Hence, we follow prior works that generate data on the GTA game engine to make GTA-Human publically available.