Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong Wang
Introduction
Achieving the ability to predict a person’s intent, preference and future activities is one of the fundamental goals for AI systems. This is particularly useful when it comes to egocentric video data for applications such as augmented reality (AR) and robotics. Imagining with an egocentric view inside the kitchen (e.g., Figure 1), if an AI system can forecast what the human would do next, an AR headset could provide useful and timely guidance, and a robot can react and collaborate with the human more smoothly.
What space should the model predict on? Recent approaches have been proposed to predict the discrete future action category given a sequence of frames as inputs, namely action anticipation. However, predicting a semantic label does not reveal how the human moves and what the human will interact with in the future. On the other hand, predicting pixels for future frames is very challenging due to its high dimension outputs with large uncertainties. Instead of adopting these two representations, our work is inspired by recent work on human motion trajectory prediction which takes images as inputs and outputs the coordinates of future pose joints. Trajectory not only provides a concrete description of motion, but also is a much smaller space to predict compared to pixel prediction. However, unlike previous works, prediction in egocentric videos also involves dense interactions with objects, which cannot be modeled by trajectory alone.
In this paper, we propose to jointly predict the future hand motion trajectory and the interaction hotspots (affordance) of the next-active object, given a sequence of input frames from an egocentric video. Starting from the final frame of the input video, we will predict the trajectories for both hands by sampling from a probabilistic distribution inferred by the model. Instead of learning a deterministic model, we tackle the uncertainty of future in a probabilistic manner. At the same time, we will predict the contact points on the next-active object interacted by the future hands. These contact points are also represented via probabilistic distributions in the form of interaction hotspots and conditioned on the predicted hand trajectory. To perform joint predictions, we introduce a Transformer-based model and an automatic way to generate a large-scale dataset for training.
Instead of collecting annotations for hand trajectories and interaction hotspots with high-cost human labor, we propose an automatic manner to collect the data in a large-scale. Given a video, we call the input frames to our model the observation frames and the predicted ones are called future frames. We first utilize off-the-shelf hand detectors to locate hands in all the future frames. Since the camera is usually moving in egocentric videos, we leverage homography in nearby frames and project the detected future hands’ locations back to the last observation frame. In this way, all the detections are aligned in the same coordinate system. Similarly, we also detect the locations where the hand interacts with the object in future frames, and project them back to the last observation frame. This process prepares the data for training our prediction model, and we generate labels for Epic-Kitchens-55, Epic-Kitchens-100 and EGTEA Gaze+ datasets without any human labor.
With the collected data, we propose to learn a novel Object-Centric Transformer (OCT) model which captures the hand-object relations from videos for hand trajectory and interaction hotspots prediction. Given the observation frames as inputs, we first extract their visual representations with a ConvNet. We perform hand and object detection and adopt RoI Align to extract their features. We take both hand and object features as object-centric tokens, and the average-pooled frame feature as image context tokens. We forward all tokens from all input frames to a Transformer encoder, which performs hand, object and environment context interaction reasoning using self-attention. Instead of decoding in a deterministic manner, we adopt the Conditional Variational Autoencoders (C-VAE) as network head in the Transformer decoder to model the uncertainty in prediction. Specifically, we compute cross-attention between the output tokens from the Transformer encoder and predicted future hand locations in the Transformer decoder. The obtained tokens are taken as conditional variables for the C-VAE. The decoder will predict both the hand trajectories and interaction hotspots jointly, and the training is supervised by a reconstruction loss corresponding to the ground-truths.
We perform evaluation on Epic-Kitchens-55 , Epic-Kitchens-100 and EGTEA Gaze+ datasets. We manually annotate the validation sets with trajectory and hotspots labels using the Amazon Mechanical Turk platform. Our OCT model significantly outperforms the baselines on both hand trajectory and interaction hotspots prediction tasks. Interestingly, we find that trajectory estimation helps interaction hotspots prediction and with more automatic annotated training data we can get better results. Finally, we experiment with fine-tuning the trained model on the action anticipation task, and find that predicting hand trajectory and interaction hotspots can benefit classifying future actions.
We propose to jointly predict hand trajectory and interaction hotspots from egocentric videos, and collect new training and test annotations.
A novel Object-Centric Transformer which models the hand and object interactions for predicting future trajectory and affordance.
We not only achieve state-of-the-art performance on both prediction tasks on Epic-Kitchens and EGTEA Gaze+ datasets, but also show our model can help the action anticipation task.
Related Work
Video anticipation. Video anticipation aims to forecast future events in videos, including future frames prediction , action anticipation , and dynamics learning . However, most of these works either relied on anticipating high-dimensional visual representations of the future, which is extremely challenging in dynamic scenes with appearance changes and moving agents, or focused on predicting a semantic label of future actions. The labels could neither tell us where the person intends to move nor the object the person would like to interact with. On the contrary, we predict the future hand motion trajectory and the interaction hotspots, both reflecting human intention and future interactions in low dimensions.
Human motion forecasting. Predicting future human motions or trajectories has been a long-standing research topic. Many of them operate on third-person vision or fixed bird’s eye view settings. Given that the first-person vision could better capture people’s intention and interactions, as well as its applicability to AR and robotics , estimating human motions in egocentric videos worth more attention. As hands are central means for humans to explore and manipulate in egocentric videos, forecasting where human hands move could reveal future activity and understand a person’s intention. Liu et al. also studied future hand trajectory estimation in egocentric videos, but their method is limited by manual annotation and single hand prediction. In contrast, we design an automatic way to collect the data on a large scale and can learn future trajectories for both hands from the data.
Grounded affordance prediction. Object affordance grounding refers to locating where the interaction occurs on an object. Given the video input, the affordance prediction task is to estimate the future active regions on an object that the human would interact. In general, there are two main categories of prediction, next-active object and interaction hotspots . The former one segments the object that will next come into contact with the hand holistically, neglecting the fine-grained spatial regions on the object’s surface. The latter one outputs a heatmap to indicate salient regions on the object. Nagarajan et al. proposed a weakly-supervised method to ground interaction hotspots on inactive images. Going one step further, we consider predicting interaction hotspots in egocentric videos. The task is more challenging since it needs to localize the next-active-object in a cluttered scene before grounding the hotspots. In our work, instead of using heatmap representation, we directly predict the contact locations more compactly.
Transformers for video forecasting. Following the immense success of Transformer in natural language processing, recent studies showed its effectiveness in solving vision tasks . The long-range reasoning and sequence modeling capability make Transformers suitable for video understanding . The TimeSformer viewed the video as a sequence of patches and adopted divided space-time attention to capture spatial-temporal relations in videos. Transformers have also been widely used in video forecasting problems like action anticipation , trajectory estimation and human motion prediction . Recent works show promising results by incorporating VAE into Transformers for generative modeling . In our work, we propose an Object-Centric Transformer (OCT) that takes the RoIAlign hand, object, and environment feature vectors extracted from a pre-trained ConvNet as input tokens. The Transformer encoder adopts self-attention across all input tokens while the Transformer decoder computes cross-attention between output tokens from the encoder and predicted future hand locations. We also introduce the C-VAE head in the Transformer decoder to express the uncertainty of the future.
Problem Setup
Given observation key frames of length as input, where is the last observation frame, our goal is to predict future hand trajectories and object contact points in the future time horizons of , as shown in Figure 2. denotes the future hand trajectories. At each time step , the future hand location consists of the left and right hand 2D pixel locations in the last observation frame. Time step is when the hand-object contact occurs and frame is the contact frame. denotes the future object contact points, where is the maximum number of contact predictions and each element defines the 2D future contact location in the last observation frame.
2 Training Data Generation
We describe how to collect training labels of future hand trajectory and object hotspots from future key frames automatically without manual labor. We first run an off-the-shelf active hand-object detector to get hand and object bounding boxes per frame, providing the future hand locations in each frame. Then we project them back to the last observation frame to collect a complete future hand trajectory. See Figure 2 for illustration. As shown in , the global motion between two consecutive frames is usually small, and they can be related by a homography . Given the homography between every two consecutive frames, we could build a chain and establish the relations of each future frame w.r.t. the last observation frame and project the future hand location back. In order to estimate the homography, we first exclude the moving objects, in particular the detected hands and objects in each frame. We mask out the corresponding location and find the correspondences between two frames outside the masked regions using SURF descriptor . We calculate the homography by sampling 4 points and applying RANSAC to maximize the number of inliers.
Similarly, for collecting the hotspots labels, we perform an additional skin segmentation and fingertip detection https://www.computervision.zone/courses/finger-counter/ within the active intersection region of hand and object bounding boxes to obtain contact points. Then we adopt a similar technique as above to project the sampled contact points from the contact frame to the last observation frame. More detailed discussions are provided in the supplementary.
3 Pre-processing
Object-Centric Transformer
is the dimension of the attention module. The attention operator computes a weighted sum of value where the weight is computed by the taking dot-product between query and key and adding up the mask followed by scaling and Softmax normalization. masks out the padding values in key by setting corresponding positions to -inf before the Softmax calculation.
The Encoder stacks multiple encoding blocks and generate outputs from inputs (Sec. 3.3):
2 Decoder
The decoder predicts future hand feature one at a time, where is the future time step. The predicted features are then sent to the trajectory head network to predict the future hand location . At each step, the decoder is auto-regressive , consuming the previously generated future hand locations as additional input when generating . The -th input to the decoder is the hand location in the last observation frame. The prediction at future time step of the decoder can be written as follows:
3 Head Networks
We employ two C-VAE as two heads; one for hand trajectory estimation and another for object contact points prediction. A C-VAE contains two functions: an encoding function which encodes the input and condition into a latent z-space parameterized by mean and co-variance , and a decoding function which decodes sampled from the latent space and condition to reconstruct input . Formally, we have and where . The and are implemented as a MLP. At training time, we minimize the objective of reconstruction error between ground-truth and predicted , as well as a KL-Divergence term that regularizes the latent z-space close to normal distribution . During inference, we sample from the latent space and concatenate with condition to predict output .
At future time step , the hand C-VAE takes hand locations as input, and conditioned on the hand feature tokens (Sec. 4.2) from the decoder output. The encoding function outputs distribution parameters and of the latent space. The decoding function predicts future hand locations . Thus, the loss function of hand C-VAE is the reconstruction loss over all future time steps and KL-Divergence regularization:
Object C-VAE.
The object C-VAE takes the future contact points sampled from the generated ground-truth set of future contact points (Sec 3.1) as input, and is conditioned on global feature token (Sec. 4.1) in the last observation frame from the Transformer encoder output and future hand locations . The hand trajectory is forwarded to a fully-connected layer and concatenated with as the conditional input. We found that the object’s future contact points could be predicted more accurately with the future hand trajectory as conditional input. During training, we use teacher forcing by taking ground-truth future hand trajectory as input. During inference, we use the predicted future hand trajectory as input for the object C-VAE. Similar to hand C-VAE, the encoding function outputs and , while the decoding function predicts future object contact points . The loss function, , of the object C-VAE is the following:
4 Training and Inference
We train the Object-Centric Transformer with both the hand trajectory loss and object contact point loss . We observe that the object contact points labels are noisier than the hand trajectory labels in our generated training set. The total loss is , where , is a constant coefficient to balance the training loss.
Inference.
During inference, we sample times for both the trajectories and contact points from the C-VAE for each input video. Following the evaluation protocol in previous work that involves stochastic unit in trajectory estimation, we report the minimum of among samples for trajectory evaluation. We collect all predicted contact points and convert them into a heatmap by centering a Gaussian distribution over each point for affordance evaluation.
Experiments
We sample frames at FPS (frames per second) as input observations and forecast s in the future on Epic-Kitchens, where the future time horizon . We sample frames at FPS on EGTEA Gaze+, forecasting s with . We use the pre-trained TSN from as the backbone to extract RGB features from the input video clip. We use the detector proposed in to detect active hand and object bounding boxes in each input frame. Then we use RoIAlign and average pooling to produce a -D vector for hand , object and global features (Sec. 3.3) at input time step . We set the embedding dimension of the OCT to . We set the number of blocks in encoder and decoder to be and on Epic-Kitchens, and on EGTEA Gaze+. Each block has attention heads. For encoding and decoding function and in C-VAE, we use a single-layer MLP for both the hand and object. The OCT is trained using Adam optimizer with a learning rate of and a batch size of . Training takes epochs on Epic-Kitchens, epochs on EGTEA Gaze+, including epochs warm-up and rest epochs with cosine decay . During inference, we sample times from the C-VAE for both hand trajectory and object contact points. Please see supplementary for detailed network structures.
2 Datasets
We use Epic-Kitchens-55 (EK55) , Epic-Kitchens-100 (EK100) and EGTEA Gaze+ (EG) datasets for experiments. The EK100 dataset is an extended version of the EK55 dataset. All datasets capture daily activities in the kitchen. Following the standard partition protocol in , we split the training set of both datasets into training and validation splits. Given the test set are only used for action anticipation, we don’t incorporate them in our experiments. We used the method in Sec. 3.2 to generate training labels automatically. The evaluation is performed on the validation split of all datasets. We manually filtered out badly generated hand trajectories and collected interaction hotspot annotations on a challenging subset via the Amazon Mechanical Turk platform (see supplementary for details). Given the last observation frame and contact frame in the future, we ask workers to place 1-5 future contact points in the last observation frame. Following , we convert these annotations into an affordance heatmap as our ground-truth. On the EK55 dataset, we collect training samples, evaluation hand trajectories, and interaction evaluation hotspots. On the EK100 dataset, we collect training samples, evaluation hand trajectories, and evaluation interaction hotspots. On the EG dataset, we collect training samples, evaluation hand trajectories, and evaluation interaction hotspots.
3 Evaluation Metrics
We use normalized predicted 2D hand locations for evaluation using the following metrics.
Interaction hotspots evaluation.
We downsample and normalize the affordance heatmap with a resolution of and ensure it sums up to . We don’t use KLD (Kullback-Leibler Divergence) metric as it is known to be sensitive to the tail of the distributions . A small difference in the low-density regions may induce a huge KLD, especially severe for forecasting problems.
Similarity Metric (SIM): SIM measures the similarity between the predicted affordance map distribution and the ground-truth one. It is computed as the sum of the minimum values at each pixel location between the predicted map and the ground-truth map.
AUC-Judd (AUC-J): AUC-J is a variant of AUC proposed by Judd et al. . The AUC evaluates the ratio of ground-truth captured by the predicted affordance map under different thresholds .
Normalized Scanpath Saliency (NSS): NSS measures the correspondence between the predicted affordance map and the ground truth. It is computed by normalizing the predicted affordance map to have zero mean and unit standard deviation and averaging over ground truth locations.
4 Comparison to the state-of-the-art
We evaluate our method against several baselines and state-of-the-art approaches. Kalman Filter (KF) tracks the center of the hand in observation frames and predicts future hands locations. Seq2Seq used LSTM to encode temporal information in the observation sequence and decode the target locations. Forecasting HOI (FHOI) used I3D (CNN) with motor attention to forecast future hand motion. Note that FHOI only used observation frames as input without accessing to hand-object detections. Besides, we also compare against Divided Attention (Divided) Transformer design by applying temporal attention and spatial attention separately in the encoder of the OCT instead of doing them jointly (Sec. 4.1). We compute temporal attention only across hand tokens in different frames and spatial attention only within each frame. The results are shown in Table 1. Experimental results show that our method outperforms previous approaches by a large margin, improving the ADE by , and FDE by on the EK100 dataset against the second-best method of each metric, and achieves similar performance with the Divided Attention Transformer encoder design. This demonstrates the superiority of using Transformer to capture hand, object, and environment context interactions in egocentric videos.
Interaction hotspots prediction.
We compare our results with the following methods. Center generated the heatmap by placing a fixed Gaussian at the center of the image. Hotspots anticipated spatial interaction regions using Grad-Cam , given the future action label as additional input. FHOI and Divided are the same method and baseline introduced in trajectory estimation, where they used I3D (CNN) and divided space-time Transformer encoder respectively. Table 2 summarizes the results of interaction hotspots prediction. Our method achieves the best performance across datasets and all metrics, improves SIM by , AUC-J by , and NSS by on the EK100 dataset against the second-best method of each metric. Compared to Divided Attention, jointly modeling all hand-object tokens in observation frames is more beneficial for the prediction. These results also highlight that the Transformer architecture is more suitable for visual forecasting problems.
Cross dataset generalization.
We evaluate learned models’ cross-dataset generalization ability on both tasks. All models are trained on Epic-Kitchens and tested on EGTEA Gaze+. The hand trajectory estimation and interaction hotspots prediction performances are shown in Table 4 and Table 4 respectively. In addition to superior in-domain performance, our method demonstrates strong cross-domain generalization by significantly outperforming other approaches across all metrics on both tasks.
5 Ablations and Analysis
We do ablation studies of our method on the EK100 dataset.
First, we evaluate the performance of using different stochastic/deterministic head networks for trajectory estimation and contact points prediction. For trajectory estimation, we compare proposed C-VAE with MLP and Bivariate. MLP deterministically outputs the future hand locations, while the Bivariate assumes the future hand location follows a bivariate Gaussian distribution at each time step and explicitly samples from the predicted distribution during inference. For future contact points prediction, we compare C-VAE with MLP and MDN. MDN adopts the Mixture Density Model (MDN) and models the distribution of future contact points as a mixture of Gaussians, where we set the number of Gaussian components to be . As shown in Table 6 and Table 6, stochastic models outperform the deterministic one on both tasks, thanks to their ability to deal with uncertainty. Adopting C-VAE against MLP improves the trajectory estimation performance by on ADE and on FDE, also obtains , , and gain on SIM, AUC-J, and NCC of hotspots prediction. Besides, we also observe that C-VAE achieves better results compared to Bivariate and MDN. It demonstrates modeling stochastic in latent space works better than output space.
C-VAE condition.
Besides modeling uncertainty in C-VAE, we analyze the effect of condition dependency in C-VAE. In Table 7, we evaluate the performance of using different C-VAE conditions for both the hand and the object. We compare three cases: no condition between the hand trajectory and object contact point, denoted as None; hand trajectory is conditioned on object contact point, denoted as ; object contact points is conditioned on hand trajectory, denoted as . We find that explicitly incorporating the conditional dependency in C-VAE improves the overall performance. Predicting interaction hotspots conditioned on future hand trajectory leads to the best result on both tasks, obtaining , , and performance gain on SIM, AUC-J, and NCC against conditioned on the inverse order. It suggests that the two tasks are intertwined and modeling their relation explicitly benefits the performance.
More training data.
As we generate our training data automatically without manual labeling, we are interested in understanding whether leveraging more automatically annotated training data can help boost performance. We trained two models under the same setting on EK55 and EK100 training split respectively. We evaluate their performances on the manually-collected EK100 validation split having no overlap with both EK55 and EK100 training splits. As shown in Table 8, we observe that a model trained with larger data (EK100) outperforms a model trained on EK55 on both tasks. This demonstrates the effectiveness of our method. Even though there is inevitable noise introduced during training data generation, our method could still learn useful representations for forecasting and it benefits from utilizing more training data. It also indicates a great potential for deploying our method on larger-scale egocentric videos.
Input ablation.
We evaluate the contributions of different input settings to the performance of trajectory estimation. Contact points prediction performance is conditioned on the trajectory, thus the input relation is not as straightforward as trajectory estimation. We evaluate the contribution of different features by removing them from the input and seeing the performance drop. As shown in Table 9, global features that encode environmental context (Sec. 3.3) and hand features are more crucial to the performance. By removing global features from input, ADE metric drops . Without hand features as input, FDE metric drops . This demonstrates that the global features are as imperative as hand features to the trajectory estimation. Besides, we also analyze the effect of different observation lengths in Figure 4. We observe that the performance improves as we incorporate more observation frames as input, which also proves our model is capable of capturing useful temporal information.
Action anticipation.
So far we have shown the capability of our method on the trajectory estimation and interaction hotspots prediction. We further investigate the potential of our trained model for action anticipation task. We only use the OCT encoder and add a single MLP on top of it that takes the output global feature token in the last observation frame (Sec. 4.1) from the encoder as an input and predicts future action labels. Following previous work , we report top-5 accuracy on EK55 dataset, and top-5 recall on EK100 dataset for verb/noun/action predictions. Each action label consists of (verb, noun). We trained our model on the same training split as we used for trajectory estimation and interaction hotspots prediction and evaluated the performance on corresponding validation splits. We only used cross-entropy loss between the prediction and ground-truth action labels during training. We obtained verb/noun prediction scores by marginalizing over the action scores. We compared two training strategies: training the model from scratch, denoted as Scratch, and fine-tuning the model pre-trained on trajectory and hotspots prediction task, denoted as Fine-tune. In the Fine-tune version, we freeze the OCT encoder and only trained the added MLP. The action anticipation performance of the two methods is shown in Table 10. The Fine-tune model outperforms Scratch model by a large margin across all datasets and all metrics. Specifically, Fine-tune obtains , , performance gain on EK55 dataset, and , , performance gain on EK100 dataset. Note that the performance of our model is not fully comparable with state-of-the-art action anticipation models as we only use a subset of samples for training and evaluation. Neither do we adopt any fancy tricks, network structures, or other loss functions for action anticipation. The experimental results show that the representation learned on trajectory estimation and interaction hotspots prediction could benefit the action anticipation task. This also proves the usefulness of the two tasks and generalization to other forecasting tasks.
Qualitative visualization.
We visualize predicted future hand trajectory and interaction hotspots in Figure 5. Our method could deal with single and two hands scenarios (when hands are visible in the last observation frame) in the first two rows. Our method can also generate diverse predictions of the future in the third row. This demonstrates that our method is able to forecast the future hand-object interaction considering the future uncertainty. Please see supplementary for more visualizations.
Discussion
Conclusion. We propose to forecast future hand-object interactions in egocentric videos. We solve this task by proposing an automatic way to collect training data, and a novel Object-Centric Transformer (OCT) model that jointly predicts future hand trajectory and interaction hotspots given a sequence of observation frames as input. Through extensive experiments and ablations, we show that OCT significantly outperforms state-of-the-art approaches, and could benefit from stochastic modeling of the future and conditional dependency of trajectory and interaction hotspots into account. Furthermore, we show that our proposed method could leverage more training data to achieve better performance and easily adapt to action anticipation task. In the future, we hope to apply our method to solve more visual forecasting problems in egocentric videos with less human supervision.
Limitation and future Work. Our training dataset generation process relies on widely used off-the-shelf tools such as active hand-object detectors and skin segmentation. Thus the ground-truth annotations for training might be affected by the bias and errors from the off-the-shelf tools. In future work, we plan to incorporate self-supervision signals during training to make our model more robust to label noise.
Acknowledgements. Prof. Wang’s lab is supported, in part, by grants from NSF CCF-2112665 (TILOS).
References
Appendix A Training Labels Generation Details
We provide additional implementation details of future hand trajectory training labels generation. As we project hand locations from all future frames to the last observation frame, we need to handle the case when there are missing hand detections in future frames. We fill the gap of missing time steps by conducting Hermite spline interpolation. Such interpolation guarantees the smoothness and continuousness of the generated trajectory. We generate the future hand trajectory at FPS and sample at FPS for training.
Interaction hotspots generation.
For interaction hotspots training label generation, we detect contact points in the contact frame and project them back to the last observation frame by a similar technique as future hand trajectory generation. However, we need to handle the active object case, i.e. the object is moved by the hand in future frames, as shown in Figure B.1. To this end, we obtain a future active object trajectory similar to the hand and move the contact points to the active object’s original place where it stays still in the last observation frame after the projection.
Appendix B Evaluation Annotations Details
We use the Amazon Mechanical Turk platform to collect interaction hotspot annotations on the evaluation set. The interface is shown in Figure B.2. We provide the contact frame in the left image and the last observation frame in the right image in the layout. Users are asked to place points on the same object location in the right image touched by the hand in the left image. The green dots are the labeled points by users. We require all the placed points to be visible and touched by the hand in the left image but haven’t been touched in the right image. We collect 1-5 contact point labels for each sample in the evaluation set.
Appendix C Implementation Details
In our proposed Transformer model, we set the embedding dimension to be and use a dropout rate of for both encoding and decoding blocks. In the C-VAE head network of both hand and object, we implement it as a -layer MLP, each for the encoding function and the decoding function . In the regular training epochs, we use cosine annealed learning rate decay starting from . During inference, as we need the hand location in the last observation frame as the -th input to the decoder, we set the normalized left hand location to and right hand location to when any of them are invisible, followed . Our model is implemented with PyTorch .
On Epic-Kitchens, our model takes s observations as input and forecasts future s hand trajectory and interaction hotspots. We sample the videos at fps for training and evaluation. We train our model for epochs, including epochs warmup.
EGTEA Gaze+.
On EGTEA Gaze+, we set the anticipation time to be s following , given it has a smaller angle of view against the Epic-Kitchens dataset. Our model takes s observations as input. We sample the videos at fps for training and evaluation. We train our model for epochs, including epochs warmup.
Appendix D Network Architectures
The network architecture is illustrated in Table B.1. We utilize ROIAlign to crop the global, hand, and object features in each input frame with dimension . Then the extracted features and the detected hand and object bounding box locations are fused in the pre-processing module to get Transformer input tokens. The tokens are passed through the OCT encoder and decoder independently. The final future hand trajectory at each time step is sampled from the hand C-VAE in an auto-regressive manner. The final object contact points are similarly sampled from the object C-VAE.
Appendix E Additional Experiments
We compare our proposed Transformer model with 3D CNN, which is widely used in video understanding. We adopt the I3D with ResNet-50 as backbone for 3D CNN. On the top of the backbone output, we predict the future hand locations and contact points by two head networks, similar to the hand and object head used in OCT. The I3D is pre-trained on Kinetics dataset. We trained the 3D CNN under the same setting as we trained OCT. The performance is shown in Table B.2. Experimental results show that by utilizing Transformer architecture against 3D CNN, we could improve the performance on both tasks. The OCT improves the FDE by and ADE by for trajectory estimation, SIM by , AUC-J by , and NSS by for hotspots prediction on the EK100 dataset. This demonstrates the superiority of adopting Transformer architecture for visual forecasting.
End-to-end training.
In the main paper, we freeze the backbone TSN and only train the OCT. We compare the performance against training end-to-end by fine-tuning the backbone along with the OCT. We apply data augmentation including random flipping and color jittering during training. The performance is shown in Table B.3. As can be seen, both models achieve comparable performance on both tasks. Given that training end-to-end is more time-consuming, we freeze the backbone in our experiments to accelerate training.
Appendix F Training Labels Visualization
We visualize the automatically generated training labels on Epic-Kitchens and EGTEA Gaze+ datasets in Figure E.3 and Figure E.4. It can be seen from the figures that our method could generate high-quality training labels under different kitchen environments and different subjects.
Appendix G Qualitative Comparisons
We compare our model’s prediction of future hand trajectory and interaction hotspots against methods that achieved second-best performance in each task, as reported in Table 1 and Table 2 in the main paper.
We visualize the prediction results on the EK100 of our method and SeqSeq that utilizes LSTM for trajectory estimation. The results are shown in Figure E.5. As can be seen, our method’s prediction is more close to the ground-truth against Seq2Seq and better reflects human intention.
Object interaction hotspots comparison.
We compare our model’s prediction of interaction hotspots against Hotspots that employ Grad-Cam to infer future hotspots map. Note that Hotspots takes the ground-truth future action label and last observation frame as input. The results are shown in Figure E.6. We observe that Hotspots’s prediction is struggling when there are multiple objects present in a cluttered scene. This implies forecasting future interaction hotspots is more challenging than the video affordance grounding task solved by Hotspots. The future hotspots estimation needs observation frames as context to locate future hand-object interactions.
Appendix H Generalization Results Visualization
We visualize our model’s prediction on the unseen environment on the EK100 dataset. The selected samples come from the validation split that contains unseen kitchens and participants. We show different samples in Figure E.7. Though the kitchen environment is unseen in training, our model could still predict reasonable future hand trajectory and interaction hotspots, which shows the in-domain generalization ability of our model.
Cross-dataset generalization.
We visualize the cross-dataset generalization ability on the EGTEA Gaze+ dataset. The model is trained on Epic-Kitchens and tested on EGTEA Gaze+. We show different samples in Figure E.8. Our model could well capture the human intention under unseen environments and subjects, forecasting future hand trajectory and interaction hotspots close to the ground-truth. It demonstrates the strong cross-domain generalization ability of our model.