Decoupling Human and Camera Motion from Videos in the Wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, Angjoo Kanazawa
Introduction
Consider the video sequence in Figure 1. As human observers, we can clearly perceive that the camera is following the athlete as he runs across the field. However, when this dynamic 3D scene is projected onto 2D images, because the camera tracks the athlete, the athlete appears to be at the center of the camera frame throughout the sequence — i.e. the projection only captures the net motion of the underlying human and camera trajectory. Thus, if we rely only on the person’s 2D motion, as many human video reconstruction methods do, we cannot recover their original trajectory in the world (Figure 1 bottom left). To recover the person’s 3D motion in the world (Figure 1 bottom right), we must also reason about how much the camera is moving.
We present an approach that models the camera motion to recover the 3D human motion in the world from videos in the wild. Our system can handle multiple people and reconstructs their motion in the same world coordinate frame, enabling us to capture their spatial relationships. Recovering the underlying human motion and their spatial relationships is a key step towards understanding humans from in-the-wild videos. Tasks, such as autonomous planning in environments with humans , or recognition of human interactions with the environment and other people , rely on information about global human trajectories. Current methods that recover global trajectories either require additional sensors, e.g. multiple cameras or depth sensors , or dense 3D reconstruction of the environment , both of which are only realistic in active or controlled capture settings. Our method acquires these global trajectories from videos in the wild, with no constraints on the capture setup, camera motion, or prior knowledge of the environment. Being able to do this from dynamic cameras is particularly relevant with the emergence of large-scale egocentric video datasets .
To do this, given an input RGB video, we first estimate the relative camera motion between frames from the static scene’s pixel motion with a SLAM system . At the same time, we estimate the identities and body poses of all detected people with a 3D human tracking system . We use these estimates to initialize the trajectories of the humans and cameras in the shared world frame. We then optimize these global trajectories over multiple stages to be consistent with both the 2D observations in the video and learned priors about how human move in the world . We illustrate our pipeline in Figure 2. Unlike existing works , we optimize over human and camera trajectories in the world frame without requiring an accurate 3D reconstruction of the static scene. Because of this, our method operates on videos captured in the wild, a challenging domain for prior methods that require good 3D geometry, since these videos rarely contain camera viewpoints with sufficient baselines for reliable scene reconstruction.
We combine two main insights to enable this optimization. First, even when the scene parallax is insufficient for accurate scene reconstruction, it still allows reasonable estimates of camera motion up to an arbitrary scale factor. In fact, in Figure 2, the recovered scene structure for the input video is a degenerate flat plane, but the relative camera motion still explains the scene parallax between frames. Second, human bodies can move realistically in the world in a small range of ways. Learned priors capture this space of realistic human motion well. We use these insights to parameterize the camera trajectory to be both consistent with the scene parallax and the 2D reprojection of realistic human trajectories in the world. Specifically, we optimize over the scale of camera displacement, using the relative camera estimates, to be consistent with the human displacement. Moreover, when multiple people are present in a video, as is often the case in in-the-wild videos, the motions of all the people further constrains the camera scale, allowing our method to operate on complex videos of people.
We evaluate our approach on EgoBody , a new dataset of videos captured with a dynamic (ego-centric) camera with ground truth 3D global human motion trajectory. Our approach achieves significant improvement upon the state-of-the-art method that also tries to recover the human motion without considering the signal provided by the background pixels . We further evaluate our approach on PoseTrack , a challenging in-the-wild video dataset originally designed for tracking. To demonstrate the robustness of our approach, we provide the results on all PoseTrack validation sequences on our project page. On top of qualitative evaluation, since there are no 3D ground-truth labels in PoseTrack, we test our approach through an evaluation on the downstream application of tracking. We show that the recovered scaled camera motion trajectory can be directly used in the PHALP system to improve tracking. The scaled camera enables more persistent 3D human registration in the 3D world, which reduces the re-identification mistakes. We provide video results and code at the project page.
Related Work
In the literature for 3D human mesh reconstruction, most methods operate by recovering the parameters of a parametric human body model, notably SMPL or its follow-up models . The main paradigms are optimization-based, e.g., SMPLify and follow-ups , or regression-based, like HMR and follow-ups . For regression approaches in particular, many efforts have focused on increasing the model robustness in a variety of settings . Most of these approaches predict the human mesh in the camera coordinate frame with identity camera. There are recent works, e.g., SPEC and CLIFF , that also consider incorporating camera information in the regression pipeline, but only for single frame inference. PHALP is a state-of-the-art method on tracking using the predicted 3D information of people ran on each frame. We use the detected identities and predicted 3D mesh as the initialization and show how it can be improved by incorporating the camera obtained by our approach.
Human Mesh Recovery from Video.
Many works extend human mesh recovery approaches on video to recover a smooth plausible human motion. However, these works fail to account for camera motion and do not recover global human trajectories. Regression approaches like HMMR , VIBE , and follow-ups operate on a bounding box level and only consider the local motion of the person within that bounding box. These approaches are prone to jitter since they are sensitive to the bounding box size. More recently, approaches such as GLAMR , D&D and Yu et al. , have tried to circumvent the issue of camera motion by recovering plausible global trajectories from the per-frame local human poses. However, relying only on local pose is not sufficient for a faithful global trajectory, especially for out-of-distribution poses, and is brittle when local pose cannot be fully observed. As such, struggles on in-the-wild videos, which often have partial occlusions and diverse human actions. Our work explicitly accounts for the camera motion to place the humans in the static scene.
Optimization approaches are similarly limited by the lack of camera awareness. use body pose smoothness priors to recover net human motion over short sequences, ignoring cameras entirely. Recent methods achieve more realistic human motion by modeling human dynamics, through learned priors or physics based priors . These priors are naturally defined in the human coordinate frame, and have thus far been limited to settings where the camera is metrically known, or static. Our approach opens a path in which these methods can be applied to moving cameras.
Other works rely on prior 3D scene information or additional sensors to contextualize human motion. can recover faithful global trajectories when the cameras and dense 3D environment have already been reconstructed. Such reconstructions require observations of the scene from many viewpoints with wide baselines. both rely on reconstructions from actively controlled capture data; rely on television data in which the same set was observed from many different viewpoints. recovers global human trajectories with multiple synchronized cameras, again only realistic for controlled capture settings, or a single static camera. In contrast, our work recovers human trajectories for in-the-wild videos, in which camera motion is uncontrolled, and the scene reconstruction is limited or non-existent. operate on monocular sequences, but the extent of results is limited to a single unoccluded person slowly walking in an indoor studio. We demonstrate our approach on PoseTrack, a complex in-the-wild dataset, which includes videos with a large number of people in various environments.
Human Mesh Recovery for Multiple People.
There have been many works that consider the reconstruction of multiple people from single images. Zanfir et al. propose an optimization approach, while follow-up work has considered regression solutions. Jiang et al. incorporate constraints that encourage the consistency of the multiple people in 3D using a Mask R-CNN type of network, while Sun et al. has investigated center-based regression . Mustafa et al. consider implicit representations for the multiple person recovery. However, all of the above works operate on a single frame basis. Mehta et al. operate on video but they only reconstruct the 3D skeleton and demonstrate results on simpler sequences with a static camera. In contrast we recover the 3D trajectories of multiple people from a moving camera.
Method
We begin by estimating each person’s per-frame pose and computing their unique identity track associations over all frames using state-of-the-art 3D tracking system, PHALP . PHALP estimates poses independently per-frame, and each estimate resides in the camera coordinate frame. In a video, however, a person’s motion in the camera coordinates is a composition of the human and camera motion in the world frame, i.e., the net motion:
To recover the original world trajectory of each person, we must determine the camera motion contribution to their net perceived motion. We denote the pose in the camera frame as , and the pose estimate in the world frame as ; the local pose and shape parameters are the same in both.
Our first insight is to use the information in the static scene’s pixel motion to compute the relative camera motion between video frames. We use state-of-the-art data-driven SLAM system, DROID-SLAM to estimate the world-to-camera transform at each time , . The camera motion can only be estimated up to an unknown scale of the world, but human bodies and motion can only take on a plausible range of values in the world. In order to ultimately place the people in the world, we must therefore determine , the relative scale between the displacement of the camera and that of people.
Our second insight is to use priors about human motion in the world to jointly determine the camera scale and people’s global trajectories. In the following sections, we describe the steps we take to initialize and prepare for joint optimization with a data-driven human motion prior. In Section 3.1, we describe how we initialize the multiple people tracks and cameras in the world coordinate frame. In Section 3.2, we describe a smoothing step on the trajectories in the world, to warm-start our joint optimization problem. Finally in Section 3.3, we describe the full optimization of trajectories and camera scale using the human motion prior.
We take as input to our joint optimization problem the pose parameters predicted by PHALP in the camera coordinate frame, , and the world-to-camera transforms estimated with SLAM, . We initialize optimization variables , for all people and timesteps . The shape and pose parameters are defined in the human canonical frame, so we use PHALP estimates directly. We initialize the global orientation and root translation in the world coordinate frame using the estimated camera transforms and camera-frame pose parameters.
where we initialize the camera scale . The joints in the world frame are then expressed as:
We use the image observations, that is, the detected 2D keypoints and confidences , to define the joint reprojection loss:
In the first stage of optimization, we align the parameters of the people in the world with the observed 2D keypoints. Because the reprojection loss (6) is very under-constrained, in this stage, we optimize only the global orientation and root translation of the human pose parameters:
We optimize Equation 7 for 30 iterations with .
2 Smoothing trajectories in the world
We next begin optimizing for the camera scale and the human shape and body pose parameters. As we begin to update , we must disambiguate the contribution of camera motion from the contribution of human translation to the reprojection error of the joints in Equation 6. To do this, we introduce additional priors about how humans move in the world to constrain the displacement of the people to be plausible. We ultimately use an data-driven transition-based human motion prior in our final stage of optimization; to prepare for this, we perform an optimization stage to smooth the transitions between poses in the world trajectories. We use a simple prior of joint smoothness, or minimal kinematic motion:
We optimize for 60 iterations and use .
3 Incorporating learned human motion priors
We finally introduce a learned motion prior that better captures the distribution of plausible human motions. We use the transition-based motion prior, HuMoR , in which the likelihood of a trajectory can be factorized into the likelihoods of transitions between consecutive states, , where is an augmented state representation used by , containing the SMPL pose parameters , as well as additional velocity and joint location predictions. The likelihood of a transition is modeled by a conditional variational autoencoder (cVAE) as
We perform optimization over the initial states , the camera scale , and latent variables , for timesteps and people . We initialize the transition latents from consecutive states and with the pre-trained HuMoR encoder , and use the HuMoR decoder to recursively roll out state from the previous state and current latent :
We recover the entire trajectories for all people by autoregressively rolling out the initial states with the latents initialized in Eq. 11. We also carry over additional losses from to regularize the predicted velocity and joint location components of to be physically plausible and consistent with the pose parameter components of ; please see for more details. We denote all prior optimization terms as
while also encouraging their distance from the ground to be less than a threshold :
Our optimization problem for this stage is then
We optimize Equation 14 with an incrementally increasing horizon, increasing in chunks of 10: , . We optimize adaptively, rolling out the trajectory by 10 more frames each time the loss decreases less than a threshold , for a minimum of 5 iterations and maximum 20 iterations. We use and . We perform all optimization with PyTorch using the L-BFGS algorithm with learning rate 1.
4 Implementation details
Handling multiple people in-the-wild: Although simple in concept, reasoning about multiple people at once in the already big optimization problem is a challenge, particularly since in videos in-the-wild, not all people appear at the same timestamp. People can enter the video at any frame, leave and come back again. Our implementation is designed to handle these cases well. We also use an improved version of PHALP with a stronger detector , which we refer to as PHALP+. Please see the appendix for more details. Code is available at the project page.
Experimental Results
We demonstrate quantitatively and qualitatively that our approach effectively reconstructs human trajectories in the world. We also demonstrate quantitatively that the camera scale we recover can be used to improve people tracking in videos. We encourage viewing additional video results on the project page.
Datasets. Datasets typically used for evaluation in the 3D human pose literature generally only provide videos captured with a static camera (e.g., Human3.6M , MPI-INF-3DHP , MuPoTS-3D , PROX ). 3DPW is a dataset captured with moving cameras, and includes indoor and outdoor sequences of people in natural environments. However, as is also discussed by previous work , it only provides 3D pose ground truth in the local frame of the person, and it is not possible to evaluate the global motion of the person. We use only 3DPW to perform ablations on the reconstructed local pose using our method.
The most relevant dataset that is captured with dynamic cameras and provides ground truth 3D pose in the global frame is the recently introduced EgoBody dataset . EgoBody is captured with a head-mounted camera on an interactor, who sees and interacts with a second interactee. The camera moves as the interactor moves, and the ground truth 3D poses of the interactee are recorded in the world frame. Because videos are recorded from a head-mounted camera, Egobody videos have heavy body truncations, with the interactee often only visible from chest or waist up.
We also demonstrate our approach on the PoseTrack dataset . PoseTrack is an extremely challenging in-the-wild dataset originally designed for tracking. It spans a wide variety of activities, involving many people with heavy occlusion and interaction. We use PoseTrack to qualitatively demonstrate the robustness of our method, and show many results in Figure 3 and in the Sup. Video. Because there is no 3D ground truth on PoseTrack, we perform quantitative evaluation through the downstream task of tracking. We show that reasoning about the tracks with the scaled camera trajectory recovered by our approach, can boost its state-of-the-art performance on PoseTrack.
Evaluation metrics: We report a variety of metrics, with a focus on metrics that compute the error on the world coordinate frame. World PA Trajectory - MPJPE (WA-MPJPE) reports MPJPE after aligning the entire trajectories of the prediction and ground truth with Procrustes Alignment. World PA First - MPJPE (W-MPJPE) reports the MPJPE after aligning the first frames of the prediction and the ground truth. PA-MPJPE reports the MPJPE error after aligning every frame of the prediction and the ground truth. We also report Acceleration Error computed as the difference between the magnitude of the acceleration vector at each joint. Please see the Sup. Mat.for more details on the evaluation protocol. For tracking, we report identity switch metrics; other commonly used tracking metrics measure quality of association and detection, but we use the same detection and association protocol in all baselines, so we omit those.
To demonstrate the effect of the different components of our system, we first perform an ablation study reporting results in Table 1. We use the metrics presented earlier and we discuss different settings. We start with the result of the full system and remove some key components. We report performance metrics of (i) our method without the last stage of optimization, i.e., without the motion prior and scale (“w/o last stage”), (ii) our method before optimization in the world, i.e., only the PHALP+ results with the estimated cameras, and (iii) the basic results of PHALP+ in the camera frame, without estimated cameras at all. We report metrics on the reconstructed trajectories in the world (W-MPJPE, WA-MPJPE, Acc Error) in Table 1.
For completeness, we also report the PA-MPJPE, which is common in the literature, for the same ablations on both the Egobody and 3DPW datasets. Because 3DPW annotates up to two people’s poses for the captured sequences, we evaluate two variants of our method. 3DPW† uses our full system’s pre-processing: PHALP+ to detect, track, and estimate people’s initial 3D local poses. 3DPW∗ uses each person’s ground-truth tracks, and runs PHALP+’s 3D pose estimation only. We note that 3DPW∗, i.e., using the ground truth tracks, is most similar to the current evaluation practices on 3DPW. We report the results in Table 2. We see that in Egobody, in which the subject is often heavily truncated, the prediction method PHALP+ achieves better performance. In other words, further optimization with truncated observations can reduce performance; this is a known issue for mesh recovery methods . However, in 3DPW, in which the subjects are more fully visible, our method improves upon the initial predictions from PHALP+. Ultimately, PA-MPJPE only captures local pose accuracy, and cannot describe the global attributes of the full trajectories.
We also compare our approach with a series of state-of-the-art methods for human mesh recovery in Table 3. The closest to our system is GLAMR , which also estimates 3D body reconstructions in the world frame. As we see in Table 3, we comfortably outperform GLAMR. Because GLAMR computes the world trajectory based on local pose estimates alone, it is especially sensitive to the extreme truncation in Egobodoy videos. In contrast to GLAMR, we leverage the relative camera motion to achieve significantly better reconstruction results, globally and locally.
We also compare against human mesh recovery baselines that compute the motion in the camera frame only. We include state-of-the-art baselines for a) single frame mesh regression (PHALP ), b) temporal mesh regression (VIBE ), and c) temporal mesh optimization (VIBE-opt ). Our method outperforms all baselines in world-level metrics, and is second best in local pose metrics.
2 Posetrack results
For the PoseTrack dataset, we show qualitative results in Figure 3, comparing against GLAMR and PHALP+ (an input to our method). We also provide qualitative reconstructions of all Posetrack validation videos in the supplemental video at the project page. We highly encourage seeing the qualitative results in video to appreciate the improvements in world trajectories. We see that our method recovers smoother trajectories that are more consistent with the dynamic scene in the input videos. We see that GLAMR struggles to plausibly position multiple people in the same world frame.
We also demonstrate a downstream application of our method, by providing helpful camera information to the PHALP tracking method . In brief, PHALP uses an estimate of the 3D person location in the camera frame for tracking. We posit that tracking is better done in the world coordinate frame, as it will be invariant to camera motion. We provide to PHALP the recovered camera motion from our approach, i.e., relative cameras from with the scale factor from our optimization to place the people in the world coordinate frame. We make minimal adaptation to the PHALP algorithm to demonstrate the effect of camera information, however there is a potential for even more improvement. Please see the supplemental for more details. The rest of the tracking procedure operates as in PHALP. We report the results in Table 4. We report both PHALP and PHALP+, which uses a better detection system , along with two variants using additional camera information: (i) using the cameras of , without rescaling to the scale of human motion, and (ii) also using the recovered scale from our optimization. We see that using out-of-the-box cameras from does not change performance. In contrast, using the recovered scale from our approach makes a significant improvement to the ID switch metric. This observation indicates our method recovers a more accurate scale, and shows the benefit it can have in the challenging tracking scenario. In Figure 4, we also demonstrate an example of the better behavior we achieve with our tracking.
Discussion
We propose a method for recovering human motion trajectories in the world coordinate frame from challenging videos with moving cameras. Our approach leverages relative camera estimates from scene pixel motion to optimize trajectories jointly with learned human motion priors for all people in the scene. This allows us to outperform state-of-the-art methods on the Egobody dataset and generate plausible trajectories for scenes with multiple people and challenging camera motions, as demonstrated by our experiments on the PoseTrack dataset.
While our system unlocks many new sources of human data, many problems remain to be addressed. In-the-wild videos often have ill-posed multiview geometry, such as predominantly rotational camera motion or co-linear motion between humans and cameras. Our method can recover inconsistent trajectories in these cases. Please see the supplemental video for examples. An exciting avenue for future work would be to incorporate human motion priors to also constrain and update the camera and scene reconstruction.
Acknowledgements: This research was supported by the DARPA Machine Common Sense program as well as BAIR/BDD sponsors, the Hellman Fellows Partnership,
References
Appendix
Appendix A Details of EgoBody evaluation
In Section 4.1 of the main manuscript, we present an experiment on the EgoBody dataset. Here, we provide more details about this evaluation.
We report results on the validation set of EgoBody. Regarding the estimated camera, we use DROID-SLAM with ground truth intrinsics. Regarding the person of interest, we first use PHALP+ (which is the same with out-of-the-box PHALP, but with a more robust detection system ), on each sequence. Since there may be multiple people in the frame (but the dataset provides 3D ground truth only for one main person), we then associate the inferred tracklets with the person of interest with the 3D ground-truth pose. For each detected bounding box, we run a 2D keypoint detection network . We run our method and our baselines on the detected tracklets using the same detections (bounding box, 2D keypoints) and ground-truth intrinsics. To accelerate inference, we split the original videos on sequences of 100 frames and we optimize each sequence separately. We report results using both local pose metrics, i.e., PA-MPJPE and global metrics that consider the global estimated trajectory across the whole reconstructed sequence. More specifically, we report results in two settings a) after aligning the predicted sequence with the ground truth sequence using Procrustes (World PA Trajectory - MPJPE), and b) after aligning the first frame of the predicted sequence with the first frame of the ground truth sequence using Procrustes (World PA First - MPJPE).
Appendix B Details of PoseTrack tracking experiment
In Section 4.2 of the main manuscript, we present an ablation where we leverage the estimated camera and the optimized scale for the purposes of tracking on the PoseTrack dataset . Here, we give more details about this implementation.
To make a direct comparison with PHALP , we make minimal modifications to the main algorithm. PHALP uses four cues; appearance, pose, 2D location and nearness of the person. We did not modify the appearance and the pose cues, but only applied the effect of the camera on the location cues, i.e., 2D location and nearness. More specifically, PHALP estimates the 3D location for each person detection in the camera frame, using a single-frame HMR model . Given our estimated camera for each frame (i.e., relative camera from and estimated world scale from our optimization), we first transform PHALP’s 3D location to the world frame (i.e., coordinate frame of the first video frame). Next, PHALP projects these 3D location to the image plane, keeping track of the 2D location, while also recording the depth (nearness) as a separate feature. For simplicity, we take the location of each detection in the world frame and, a) keep the part of the location of each detection to represent the 2D location, (after normalizing it to $Z$ coordinate to compute the nearness. The rest of the pipeline remains the same as PHALP. Essentially, the only difference is that the location of the people are considered in the world coordinate instead of the camera coordinate frame.
We highlight that we only make minimal adaptations to the main PHALP algorithm to demonstrate the effect of camera information for tracking, but there is further room for improvement. For example, considering that we have access to the explicit 3D location for each detection in the world frame, we could also explore tracking using 3D location as a cue, instead of splitting the position cue to 2D location and depth/nearness, but this would require modification to the PHALP’s tracking parameters. Similarly, we could leverage our optimized results to compute more reliable affinity metrics on the pose, but here our goal was to decouple the benefit of the better camera from other cues, i.e., our more stable pose. It would be an interesting direction for future work to integrate all these updates and implement a more robust tracking system using information for camera motion.
Appendix C Additional implementation details
When multiple people are on the same floor level, our optimization becomes better constrained because all of them need to share the same floor , meaning that the motion of more people provides constraints for the optimization of the variable. However, in many real world videos, people are in different floor levels. In that case, when we observe that it is not possible to solve Equation 14 with a single floor variable , we separate the people in clusters based on the locations of their feet, and introduce separate floor variables . The people in cluster shares the same floor and the optimization continues as usual.
Handling multiple people
A distinct challenge of in-the-wild videos is properly handling multiple person tracks of undetermined length as they undergo occlusion. During the first two stages of optimization, each person’s pose is optimized independently. During these stages, we only optimize the people that are visible, and mask out losses on the predictions of any frame and any track that are not visible.
During the last stage, optimizing all tracks in a single batch allows scale and ground contact information to be shared between people. To do this in our incremental optimization scheme (described in Section 3.4 of the main text), we store each track with respect to its first appearance, rather than with respect to the first frame of the video. We pad the end of each track to be , the length of the longest track. Specifically, for each track, we store the start and end times of the track, , and latent vectors . The latents of each track are contiguous in time (we infill occlusions between the first and last appearances), but do not all start or end at the same timestep.
In an optimization step at the rollout horizon , we roll out steps of each track , where is the decoded latent state. We then scatter each track into the interval of input video’s timeline. That is, each state synchronized to the original time it occurred in, and remove the padded states. We then only optimize the track over the time segment containing , , and mask out the frames of each track that fall outside of this interval.
The runtime of optimization grows linearly with the number of people we track. Optimizing a sequence of around 100 frames and 4 people requires around 40 minutes.
Appendix D Robustness
One of our observations with regards to using the HuMoR motion prior is that it can be challenging to optimize, especially over a long sequence. This results in our decision to optimize the pose sequences of every person in a rollout horizon, as described in the previous section. This increases the robustness of the optimization for longer sequences and it should be applicable to any motion prior that also models the transition, e.g., .
Moreover, HuMoR assumes static camera. When used on sequences with camera movement, without modeling the camera motion as we do, it can lead to catastrophic failures in the optimization. For example, in Egobody, we observed that HuMoR fails on 30% of the sequences when we use identity (static) camera. In contrast to that, our approach, even with imperfect camera motion, i.e., using the estimates from as we do, leads to successful optimization in 99% of the sequences; for the rare cases where optimization of the HuMoR motion prior fails, we simply revert back to the results of the previous step where we optimize with the smoothness motion prior.
On the more challenging PoseTrack sequences, we also observe some rare optimization failures. Most of those are related to the single floor assumption and can be addressed by clustering the people in different floors, as described in Section 3.4 of the main manuscript.
Appendix E Limitations
One of the limitation of our approach is that we rely on outputs from other methods (e.g., estimated camera from with approximate intrinsics for in-the-wild videos, person tracking from ), which sometime can propagate failures to our optimization.
For example, SfM approaches often have trouble distinguishing between translational and rotational motions, particularly with large focal length. Although our optimization can typically infer reasonable motions even with these imperfect camera estimation, an exciting future work is to jointly optimize the camera motion and human motion, which requires also updating the 3D structure.
Another failure mode is in case of identity switch errors in tracking, with the most harmful being errors that merge two different people into a single tracklet. Although we do not explicitly reason about tracklet identity during our optimization, we provide an experiment where PHALP makes better use of information about camera motion (main manuscript, Section 4.2). Future work could also solve the association problem while optimizing over people and camera’s motion.
Finally, we observed some inherently challenging motions to decouple from a monocular video, e.g., when people move co-linearly with the camera. In these cases, our approach can underestimate the location evolution of the people, e.g., causing people to run in the same location. Please see the example in the supplemental video. In these situations, future work could consider also priors for the background scale, e.g., by using monocular depth cues , which could help to better constrain the scale factor .