Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning

Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B. Tenenbaum, Chuang Gan

Introduction

Visual reasoning in dynamic scenes involves both the understanding of compositional properties, relationships, and events of objects, and the inference and prediction of their temporal and causal structures. As depicted in Fig. 1, to answer the question “What will happen next?” based on the observed video frames, one needs to detect the object trajectories, predict their dynamics, analyze the temporal structures, and ground visual objects and events to get the answer “The blue sphere and the yellow object collide”.

Recently, various end-to-end neural network-based approaches have been proposed for joint understanding of video and language (Lei et al., 2018; Fan et al., 2019). While these methods have shown great success in learning to recognize visually complex concepts, such as human activities (Xu et al., 2017; Ye et al., 2017), they typically fail on benchmarks that require the understanding of compositional and causal structures in the videos and text (Yi et al., 2020). Another line of research has been focusing on building modular neural networks that can represent the compositional structures in scenes and questions, such as object-centric scene structures and multi-hop reasoning (Andreas et al., 2016; Johnson et al., 2017b; Hudson & Manning, 2019). However, these methods are designed for static images and do not handle the temporal and causal structure in dynamic scenes well, leading to inferior performance on video causal reasoning benchmark CLEVRER (Yi et al., 2020).

To model the temporal and causal structures in dynamic scenes, Yi et al. (2020) proposed an oracle model to combine symbolic representation with video dynamics modeling and achieved state-of-the-art performance on CLEVRER. However, this model requires videos with dense annotations for visual attributes and physical events, which are impractical or extremely labor-intensive in real scenes. We argue that such dense explicit video annotations are unnecessary for video reasoning, since they are naturally encoded in the question answer pairs associated with the videos. For example, the question answer pair and the video in Fig. 1 can implicitly inform a model what the concepts “sphere”, “blue”, “yellow” and “collide” really mean. However, a video may contain multiple fast-moving occluded objects and complex object interactions, and the questions and answers have diverse forms. It remains an open and challenging problem to simultaneously represent objects over time, train an accurate dynamic model from raw videos, and align objects with visual properties and events for accurate temporal and causal reasoning, using vision and language as the only supervision.

Our main ideas are to factorize video perception and reasoning into several modules: object tracking, object and event concept grounding, and dynamics prediction. We first detect objects in the video, associating them into object tracks across the frames. We can then ground various object and event concepts from language, train a dynamic model on top of object tracks for future and counterfactual predictions, analyze relationships between events, and answer queries based on these extracted representations. All these modules can be trained jointly by watching videos and reading paired questions and answers.

To achieve this goal, we introduce Dynamic Concept Learner (DCL), a unified neural-symbolic framework for recognizing objects and events in videos and analyzing their temporal and causal structures, without explicit annotations on visual attributes and physical events such as collisions during training. To facilitate model training, a multi-step training paradigm has been proposed. We first run an object detector on individual frames and associate objects across frames based on a motion-based correspondence. Next, our model learns concepts about object properties, relationships, and events by reading paired questions and answers that describe or explain the events in the video. Then, we leverage the acquired visual concepts in the previous steps to refine the object association across frames. Finally, we train a dynamics prediction network (Li et al., 2019b) based on the refined object trajectories and optimize it jointly with other learning parts in this unified framework. Such a training paradigm ensures that all neural modules share the same latent space for representing concepts and they can bootstrap the learning of each other.

We evaluate DCL’s performance on CLEVRER, a video reasoning benchmark that includes descriptive, explanatory, predictive, and counterfactual reasoning with a uniform language interface. DCL achieves state-of-the-art performance on all question categories and requires no scene supervision such as object properties and collision events. To further examine the grounding accuracy and transferability of the acquired concepts, we introduce two new benchmarks for video-text retrieval and spatial-temporal grounding and localization on the CLEVRER videos, namely CLEVRER-Retrieval and CLEVRER-Grounding. Without any further training, our model generalizes well to these benchmarks, surpassing the baseline by a noticeable margin.

Related Work

Our work is related to reasoning and answering questions about visual content. Early studies like (Wu et al., 2016; Zhu et al., 2016; Gan et al., 2017) typically adopted monolithic network architectures and mainly focused on visual understanding. To perform deeper visual reasoning, neural module networks were extensively studied in recent works (Johnson et al., 2017a; Hu et al., 2018; Hudson & Manning, 2018; Amizadeh et al., 2020), where they represent symbolic operations with small neural networks and perform multi-hop reasoning. Some previous research has also attempted to learn visual concepts through visual question answering (Mao et al., 2019). However, it mainly focused on learning static concepts in images, while our DCL aims at learning dynamic concepts like moving and collision in videos and at making use of these concepts for temporal and causal reasoning.

Later, visual reasoning was extended to more complex dynamic videos (Lei et al., 2018; Fan et al., 2019; Li et al., 2020; 2019a; Huang et al., 2020). Recently, Yi et al. (2020) proposed CLEVRER, a new video reasoning benchmark for evaluating computational models’ comprehension of the causal structure behind physical object interaction. They also developed an oracle model, combining the neuro-symbolic visual question-answering model (Yi et al., 2018) with the dynamics prediction model (Li et al., 2019b), showing competitive performance. However, this model requires explicit labels for object attributes, masks, and spatio-temporal localization of events during training. Instead, our DCL has no reliance on any labels for objects and events and can learn these concepts through natural supervision (i.e., videos and question-answer pairs).

Our work is also related to temporal and relational reasoning in videos via neural networks (Wang & Gupta, 2018; Materzynska et al., 2020; Ji et al., 2020). These works typically rely on specific action annotations, while our DCL learns to ground object and event concepts and analyze their temporal relations through question answering. Recently, various benchmarks (Riochet et al., 2018; Bakhtin et al., 2019; Girdhar & Ramanan, 2020; Baradel et al., 2020; Gan et al., 2020) have been proposed to study dynamics and reasoning in physical scenes. However, these datasets mainly target at pure video understanding and do not contain natural language question answering. Much research has been studying dynamic modeling for physical scenes (Lerer et al., 2016; Battaglia et al., 2013; Mottaghi et al., 2016; Finn et al., 2016; Shao et al., 2014; Fire & Zhu, 2015; Ye et al., 2018; Li et al., 2019b). We adopt PropNet (Li et al., 2019b) for dynamics prediction and feed the predicted scenes to the video feature extractor and the neuro-symbolic executor for event prediction and question answering.

While many works (Zhou et al., 2019; 2018; Gan et al., 2015) have been studying on the problems of understanding human actions and activities (e.g., running, cooking, cleaning) in videos, our work’s primary goal is to design a unified framework for learning physical object and event concepts (e.g., collision, falling, stability). These tasks are of great importance in practical applications such as industrial robot manipulation which requires AI systems with human-like physical common sense.

Dynamic Concept Learner

In this section, we introduce a new video reasoning model, Dynamic Concept Learner (DCL), which learns to recognize video attributes, events, and dynamics and to analyze their temporal and causal structures, all through watching videos and answering corresponding questions. DCL contains five modules, 1) an object trajectory detector, 2) video feature extractor, 3) a dynamic predictor, 4) a language program parser, and 5) a neural symbolic executor. As shown in Fig. 2, given an input video, the trajectory detector detects objects in each frame and associates them into trajectories; the feature extractor then represents them as latent feature vectors. After that, DCL quantizes the objects’ static concepts (i.e., color, shape, and material) by matching the latent object features with the corresponding concept embeddings in the executor. As these static concepts are motion-independent, they can be adopted as an additional criteria to refine the object trajectories. Based on the refined trajectories, the dynamics predictor predicts the objects’ movement and interactions in future and counterfactual scenes. The language parser parses the question and choices into functional programs, which are executed by the program executor on the latent representation space to get answers.

The object and event concept embeddings and the object-centric representation share the same latent space; answering questions associated with videos can directly optimize them through backpropagation. The object trajectories and dynamics can be refined by the object static attributes predicted by DCL. Our framework enjoys the advantages of both transparency and efficiency, since it enables step-by-step investigations of the whole reasoning process and has no requirements for explicit annotations of visual attributes, events, and object masks.

Given a video, the object trajectory detector detects object proposals in each frame and connects them into object trajectories O={on}n=1NO=\{o^{n}\}_{n=1}^{N}, where on={btn}t=1To^{n}=\{b^{n}_{t}\}_{t=1}^{T} and NN is the number of objects in the video. bt=[xtn,ytn,wtn,htn]b_{t}=[x^{n}_{t},y^{n}_{t},w_{t}^{n},h_{t}^{n}] is an object proposal at frame tt and TT is the frame number, where (xtn,ytn)(x^{n}_{t},y^{n}_{t}) denotes the normalized proposal coordinate center and wtnw^{n}_{t} and htnh^{n}_{t} denote the normalized width and height, respectively.

The object detector first uses a pre-trained region proposal network (Ren et al., 2015) to generate object proposals in all frames, which are further linked across connective frames to get all objects’ trajectories. Let {bti}i=1N\{b_{t}^{i}\}_{i=1}^{N} and {bt+1j}j=1N\{b_{t+1}^{j}\}_{j=1}^{N} to be two sets of proposals in two connective frames. Inspired by Gkioxari & Malik (2015), we define a connection score sls_{l} between btib_{t}^{i} and bt+1jb_{t+1}^{j} to be

where sc(bti)s_{c}(b_{t}^{i}) is the confidence score of the proposal btib_{t}^{i}, IoUIoU is the intersection over union and λ1\lambda_{1} is a scalar. Gkioxari & Malik (2015) adopts a greedy algorithm to connect the proposals without global optimization. Instead, we assign boxes {bt+1j}j=1N\{b_{t+1}^{j}\}_{j=1}^{N} at the t+1t+1 frame to {bti}i=1N\{b_{t}^{i}\}_{i=1}^{N} by a linear sum assignment.

Video Feature Extraction.

Grounding Object and Event Concepts.

Video Reasoning requires a model to ground object and event concepts in videos. DCL achieves this by matching object and event representation with object and event embeddings in the symbolic executor. Specifically, DCL calculates the confidence score that the nn-th object is moving by [cos⁡(smoving,mda(fns))−δ]/λ\left[\cos(s^{\text{moving}},m_{da}(f^{s}_{n}))-\delta\right]/\lambda, where fnsf^{s}_{n} denotes the temporal sequence feature for the nn-th object, smovings^{\text{moving}} denotes a vector embedding for concept moving, and mdam_{da} denotes a linear transformation, mapping fnsf^{s}_{n} into the dynamic concept representation space. δ\delta and λ\lambda are the shifting and scaling scalars, and cos⁡()\cos() calculates the cosine similarity between two vectors. DCL grounds static attributes and the collision event similarly, matching average visual features and interactive features with their corresponding concept embeddings in the latent space. We give more details on the concept and event quantization in Appendix E.

Trajectory Refinement.

The connection score in Eq. 1 ensures the continuity of the detected object trajectories. However, it does not consider the objects’ visual appearance; therefore, it may fail to track the objects and may connect inconsistent objects when different objects are close to each other and moving rapidly. To detect better object trajectories and to ensure the consistency of visual appearance along the track, we add a new term to Eq. 1 and re-define the connection score to be

where fappear({bmi}m=0t,bt+1j)f_{\text{appear}}(\{b_{m}^{i}\}_{m=0}^{t},b_{t+1}^{j}) measures the attribute similarity between the newly added proposal bt+1jb_{t+1}^{j} and all proposals ({bmi}m=0t(\{b_{m}^{i}\}_{m=0}^{t} in previous frames. We define fappearf_{\text{appear}} as

where attr∈{color,material,shape}\text{attr}\in\{\text{color},\text{material},\text{shape}\}. fattr(bm,bt+1)f_{\text{attr}}(b_{m},b_{t+1}) equals to 1 when bmib_{m}^{i} and bt+1jb_{t+1}^{j} have the same attribute, and 0 otherwise. In Eq. 2, fappearf_{\text{appear}} ensures that the detected trajectories have consistent visual appearance and helps to distinguish the correct object when different objects are close to each other in the same frame. These additional static attributes, including color, material, and shape, are extracted without explicit annotation during training. Specifically, we quantize the attributes by choosing the concept whose concept embedding has the best cosine similarity with the object feature. We iteratively connect proposals at the t+1t+1 frame to proposals at the tt frame and get a set of object trajectories O={on}n=1NO=\{o^{n}\}_{n=1}^{N}, where on={btn}t=1To^{n}=\{b^{n}_{t}\}_{t=1}^{T}.

Dynamic Prediction.

Given an input video and the refined trajectories of objects, we predict the locations and RGB patches of the objects in future or counterfactual scenes with a Propagation Network (Li et al., 2019b). We then generate the predicted scenes by pasting RGB patches into the predicted locations. The generated scenes are fed to the feature extractor to extract the corresponding features. Such a design enables the question answer pairs associated with the predicted scenes to optimize the concept embeddings and requires no explicit labels for collision prediction, leading to better optimization. This is different from Yi et al. (2020), which requires dense collision event labels to train a collision classifier.

To predict the locations and RGB patches, the dynamic predictor maintains a directed graph ⟨V,D⟩=⟨{vn}n=1N,{dn1,n2}n1=1,n2=1N,N⟩\left\langle V,D\right\rangle=\left\langle\{v_{n}\}_{n=1}^{N},\{d_{n_{1},n_{2}}\}_{n_{1}=1,n_{2}=1}^{N,N}\right\rangle. The nn-th vertex vnv_{n} is represented by a concatenation of tuple ⟨btn,ptn⟩\left\langle b^{n}_{t},p^{n}_{t}\right\rangle over a small time window ww, where btn=[xtn,ytn,wtn,htn]b^{n}_{t}=[x^{n}_{t},y^{n}_{t},w_{t}^{n},h_{t}^{n}] is the nn-th object’s normalized coordinates and ptnp^{n}_{t} is a cropped RGB patch centering at (xtn,ytn)(x^{n}_{t},y^{n}_{t}). The edge dn1,n2d_{n_{1},n_{2}} denotes the relation between the n1n_{1}-th and n2n_{2}-th objects and is represented by the concatenation of the normalized coordinate difference btn1−btn2b^{n_{1}}_{t}-b^{n_{2}}_{t}. The dynamic predictor performs multi-step message passing to simulate instantaneous propagation effects.

During inference, the dynamics predictor predicts the locations and patches at frame k+1k+1 using the features of the last ww observed frames in the original video. We get the predictions at frame k+2k+2 by autoregressively feeding the predicted results at frame k+1k+1 as the input to the predictor. To get the counterfactual scenes where the nn-th object is removed, we remove the nn-th vertex and its associated edges from the input to predict counterfactual dynamics. Iteratively, we get the predicted normalized coordinates {b^k′n}n=1,k′=1N,K′\{\hat{b}_{k^{\prime}}^{n}\}_{n=1,k^{\prime}=1}^{N,K^{\prime}} and RGB patches {p^k′n}n=1,k′=1N,K\{\hat{p}_{k^{\prime}}^{n}\}_{n=1,k^{\prime}=1}^{N,K} at all predicted K′K^{\prime} frames. We give more details on the dynamic predictor at Appendix C.

Language Program Parsing.

The language program parser aims to translate the questions and choices into executable symbolic programs. Each executable program consists of a series of operations like selecting objects with certain properties, filtering events happening at a specific moment, finding the causes of an event, and eventually enabling transparent and step-by-step visual reasoning. Moreover, these operations are compositional and can be combined to represent questions with various compositionality and complexity. We adopt a seq2seq model (Bahdanau et al., 2015) with an attention mechanism to translate word sequences into a set of symbolic programs and treat questions and choices, separately. We give detailed implementation of the program parser in Appendix D.

Symbolic Execution.

Given a parsed program, the symbolic executor explicitly runs it on the latent features extracted from the observed and predicted scenes to answer the question. The executor consists of a series of functional modules to realize the operators in symbolic programs. The last operator’s output is the answer to the question. Similar to Mao et al. (2019), we represent all object states, events, and results of all operators in a probabilistic manner during training. This makes the whole execution process differential w.r.t. the latent representations from the observed and predicted scenes. It enables the optimization of the feature extractor and concept embeddings in the symbolic executor. We provide the implementation of all the operators in Appendix E.

2 Training and inference

Training. We follow a multi-step training paradigm to optimize the model: 1) We first extract object trajectories with the scoring function in Eq. 1 and optimize the video feature extractor and concept embeddings in the symbolic executor with only descriptive and explanatory questions; 2) We quantize the static attributes for all objects with the feature extractor and the concept embeddings learned in Step 1) and refine object trajectories with the scoring function Eq. 2; 3) Based on the refined trajectories, we train the dynamic predictor and predict dynamics for future and counterfactual scenes; 4) We train the full DCL with all the question answer pairs and get the final model. The program executor is fully differentiable w.r.t. the feature extractor and concept embeddings. We use cross-entropy loss to supervise open-ended questions and use mean square error loss to supervise counting questions. We provide specific loss functions for each module in Appendix H.

Inference. During inference, given an input video and a question, we first detect the object trajectories and predict their motions and interactions in future and counterfactual scenes. We then extract object and event features for both the observed and predicted scenes with the feature extractor. We parse the questions and choices into executable symbolic programs. We finally execute the programs on the latent feature space and get the answer to the question.

Experiments

To show the proposed DCL’s advantages, we conduct extensive experiments on the video reasoning benchmark CLEVRER. Existing other video datasets either ask questions about the complex visual context (Tapaswi et al., 2016; Lei et al., 2018) or study dynamics and reasoning without question answering (Girdhar & Ramanan, 2020; Baradel et al., 2020). Thus, they are unsuitable for evaluating video causal reasoning and learning object and event concepts through question answering. We first show its strong performance on video causal reasoning. Then, we show DCL’s ability on concept learning, predicting object visual attributes and events happening in videos. We show DCL’s generalization capacity to new applications, including CLEVRER-Grounding and CLEVRER-Retrieval. We finally extend DCL to a real block tower video dataset (Lerer et al., 2016).

Following the experimental setting in Yi et al. (2020), we train the language program parser with 1000 programs for all question types. We train all our models without attribute and event labels. Our models for video question answering are trained on the training set, tuned on the validation set, and evaluated in the test set. To show DCL’s generalization capacity, we build CLEVRER-Grounding and CLEVRER-Retrieval datasets from the original CLEVRER videos and their associated video annotations. We provide more implementation details in Appendix A.

2 Comparisons on Temporal and Causal Reasoning

We compare our DCL with previous methods on CLEVRER, including Memory (Fan et al., 2019), IEP (Johnson et al., 2017b), TbD-net (Mascharka et al., 2018), TVQA+ (Lei et al., 2018), NS-DR (Yi et al., 2020), MAC (V) (Hudson & Manning, 2018) and its attribute-aware variant, MAC (V+). We refer interested readers to CLEVRER (Yi et al., 2020) for more details. Additionally, we also include a recent state-of-the-art VQA model HCRN (Le et al., 2020) for performance comparison, which adopts a conditional relation network for representation and reasoning over videos. To provide more extensive analysis, we introduce DCL-Oracle by adding object attribute and collision supervisions into DCL’s training. We summarize their requirement for visual labels and language programs in the second and third columns of Table 1.

According to the results in Table 1, we have the following observations. Although HCRN achieves state-of-the-art performance on human-centric action datasets (Jang et al., 2017; Xu et al., 2017; 2016), it only performs slightly better than Memory and much worse than NS-DR on CLEVRER. We believe the reason is that HCRN mainly focuses on motion modeling across frames while CLEVRER requires models to perform dynamic visual reasoning on videos and analyze its temporal and causal structures. NS-DR performs best among all the baseline models, showing the power of combining symbolic representation with dynamics modeling. Our model achieves the state-of-the-art question answering performance on all kinds of questions even without visual attributes and event labels from simulations during training, showing its effectiveness and label-efficiency. Compared with NS-DR, our model achieves more significant gains on predictive and counterfactual questions than that on the descriptive questions. This shows DCL’s effectiveness in modeling for temporal and causal reasoning. Unlike NS-DR, which directly predicts collision event labels with its dynamic model, DCL quantizes concepts and executes symbolic programs in an end-to-end training manner, leading to better predictions for dynamic concepts. DCL-Oracle shows the upper-bound performance of the proposed model to ground physical object and event concepts through question answering.

3 Evaluation of Object and Event Concept Grounding in Videos

Previous methods like MAC (V) and TbD-net (V) did not learn explicit concepts during training, and NS-DR required intrinsic attribute and event labels as input. Instead, DCL can directly quantize video concepts, including static visual attributes (i.e. color, shape, and material), dynamic attributes (i.e. moving and stationary) and events (i.e. in, out, and collision). Specifically, DCL quantizes the concepts by mapping the latent object features into the concept space by linear transformation and calculating their cosine similarities with the concept embeddings in the neural-symbolic executor.

We predict the static attributes of each object by averaging the visual object features at each sampled frame. We regard an object to be moving if it moves at any frame, and otherwise stationary. We consider there is a collision happening between a pair of objects if they collide at any frame of the video. We get the ground-truth labels from the provided video annotation and report the accuracy in table 2 on the validation set.

We observe that DCL can learn to recognize different kinds of concepts without explicit concept labels during training. This shows DCL’s effectiveness to learn object and event concepts through natural question answering. We also find that DCL recognizes static attributes and events better than dynamic attributes. We further find that DCL may misclassify objects to be “stationary” if they are missing for most frames and only move slowly at specific frames. We suspect the reason is that we only learn the dynamic attributes through question answering and question answering pairs for such slow-moving objects rarely appear in the training set.

4 Generalization

We further apply DCL to two new applications, including CLEVRER-Grounding, spatio-temporal localization of objects or events in a video, and CLEVRER-Retrieval, finding semantic-related videos for the query expressions and vice versa.

We first build datasets for video grounding and video-text retrieval by synthesizing language expressions for videos in CLEVRER. We generate the expressions by filling visual contents from the video annotations into a set of pre-defined templates. For example, given the text template, “The <<static_attribute>> that is <<dynamic_attribute>> <<time_identifier>>”, we can fill it and generate “The metal cube that is moving when the video ends.”. Fig. 3 shows examples for the generated datasets, and we provide more statistics and examples in Appendix G. We transform the grounding and retrieval expressions into executable programs by training new language parsers on the expressions of the synthetic training set. To provide more extensive comparisons, we adopt the representative video grounding/ retrieval model WSSTG (Chen et al., 2019) as a baseline. We provide more details of the baseline implementation in Appendix A.

CLEVRER-Grounding contains object grounding and event grounding. For video object grounding, we localize each described object’s whole trajectory and compute the mean intersection over union (IoU) with the ground-truth trajectory. For event grounding, including collision, in and out, we temporally localize the frame that the event happens at and calculate the frame difference with the ground-truth frames. For collision event, we also spatially localize the collided objects’ the union box and compute it’s IoU with the ground-truth. We don’t perform spatial localization for in and out events since the target object usually appears to be too small to localize at the frame it enters or leaves the scene.

Table 4 lists the results. From the table, we can find that our proposed DCL transforms to the new CLEVRER-Grounding task well and achieves high accuracy for spatial localization and low frame differences for temporal localization. On the contrary, the traditional video grounding method WSSTG performs much worse, since it mainly aligns simple visual concepts between text and images and has difficulties in modeling temporal structures and understanding the complex logic.

CLEVRER-Retrieval.

For CLEVRER-Retrieval, an expression-video pair is considered as a positive pair if the video contains the objects and events described by the expression and otherwise negative. Given a video, we define its matching similarity with the query expression to be the maximal similarity between the query expression and all the object or event proposals. Additionally, we also introduce a recent state-of-the-art video-text retrieval model HGR (Chen et al., 2020) for performance comparison, which decomposes video-text matching into global-to-local levels and performs cross-modal matching with attention-based graph reasoning. We densely compare every possible expression-video pair and use mean average precision (mAP) as the retrieval metric.

We report the retrieval mAP in Table 4. Compared with CLEVRER-Grounding, CLEVRER-Retrieval is more challenging since it contains many more distracting objects, events and expressions. WSSTG performs worse on the retrieval setting because it does not model temporal structures and understand its logic. HGR achieves better performance than the previous baseline WSSTG since it performs hierarchical modeling for events, actions and entities. However, it performs worse than DCL since it doesn’t explicitly model the temporal structures and the complex logic behind the video-text pairs in CLEVRER-Retrieval. On the other hand, DCL is much more robust since it can explicitly ground object and event concepts, analyze their relations and perform step-by-step visual reasoning.

5 Extension to real videos and the new concept

We further conduct experiments on a real block tower video dataset (Lerer et al., 2016) to learn the new physical concept “falling”. The block tower dataset has 493 videos and each video contains a stable or falling block tower. Since the original dataset aims to study physical intuition and doesn’t contain question-answer pairs, we manually synthesize question-answer pairs in a similar way to CLEVRER (Yi et al., 2020). We show examples of the new dataset in Fig 4. We train models on randomly-selected 393 videos and their associated question-answer pairs and evaluate their performance on the rest 100 videos.

Similar to the setting in CLEVRER, we use the average visual feature from ResNet-34 for static attribute prediction and temporal sequence feature for the prediction of the new dynamic concept “falling”. Additionally, we train a visual reasoning baseline MAC (V) (Hudson & Manning, 2018) for performance comparison. Table 6 lists the results. Our model achieves better question-answering performance on the block tower dataset especially on the counting questions like “How many objects are falling?”. We believe the reason is that counting questions require a model to estimate the states of each object. MAC (V) just simply adopts an MLP classifier to predict each answer’s probability and doesn’t model the object states. Differently, DCL answers the counting questions by accumulating the probabilities of each object and is more transparent and accurate. We also show the accuracy of color and “falling” concept prediction on the validation set in Table 6. Our DCL can naturally learn to ground the new dynamic concept “falling” in the real videos through question answering. This shows DCL’s effectiveness and strong generalization capacity.

Discussion and future work

We present a unified neural symbolic framework, named Dynamic Concept Learner (DCL), to study temporal and causal reasoning in videos. DCL, learned by watching videos and reading question-answers, is able to track objects across different frames, ground physical object and event concepts, understand the causal relationship, make future and counterfactual predictions and combine all these abilities to perform temporal and causal reasoning. DCL achieves state-of-the-art performance on the video reasoning benchmark CLEVRER. Based on the learned object and event concepts, DCL generalizes well to spatial-temporal object and event grounding and video-text retrieval. We also extend DCL to real videos to learn new physical concepts.

Our DCL suggests several future research directions. First, it still requires further exploration for dynamic models with stronger long-term dynamic prediction capability to handle some counterfactual questions. Second, it will be interesting to extend our DCL to more general videos to build a stronger model for learning both physical concepts and human-centric action concepts.

Acknowledgement This work is in part supported by ONR MURI N00014-16-1-2007, the Center for Brain, Minds, and Machines (CBMM, funded by NSF STC award CCF-1231216), the Samsung Global Research Outreach (GRO) Program, Autodesk, and IBM Research.

References

Appendix A Implementation Details

Since it’s extremely computation-intensive to predict the object states and events at every frame, we evenly sample 32 frames for each video. All models are trained using Adam (Kingma & Ba, 2014) for 20 epochs and the learning rate is set to 10−410^{-4}. We adopt a two-stage training strategy for training the dynamic predictor. For the dynamic predictor, we set the time window size ww, the propogation step LL and dimension of hidden states to be 3, 2 and 512, respectively. Following the sample rate at the observed frames, we sample a frame for prediction every 4 frames. We first train the dynamic model with only location prediction and then train it with both location and RGB patch prediction. Experimentally, we find this training strategy provides a more stable prediction. We train the language parser with the same training strategy as Yi et al. (2018) for fair comparison.

Baseline Implementation.

We implement the baselines HCRN (Le et al., 2020), HGR Chen et al. (2020) and WSSTG (Chen et al., 2019) carefully based on the public source code. WSSTG first generate a set of object or event candidates and match them with the query sentence. We choose the proposal candidate with the best similarity as the grounding result. For object grounding, we use the same tube trajectory candidates as we use for implementing DCL. For grounding event concepts in and out, we treat each object at each sampled frame as a potential candidate for selection. For grounding event concept collision, we treat the union regions of any object pairs as candidates. For CLEVRER-Retrieval, we treat the proposal candidate with the best similarity as the similarity score between the video and the query sentence. We train WSSTG with a synthetic training set generated from the videos of CLEVRER-VQA training set. A fully-supervised triplet loss is adopted to optimize the model.

Appendix B Feature Extraction

We evenly sample KK frames for each video and use a ResNet-34 (He et al., 2016) to extract visual features. For the nn-th object in the video, we define its average visual feature to be fnv=1K∑k=1Kfknf^{v}_{n}=\frac{1}{K}\sum_{k=1}^{K}f_{k}^{n}, where fknf_{k}^{n} is the concatenation of the regional feature and the global context feature at the kk-th frame. We define its temporal sequence feature fnsf_{n}^{s} to be the contenation of [xtn,ytn,wtn,htn][x^{n}_{t},y^{n}_{t},w_{t}^{n},h_{t}^{n}] at all TT frames, where (xtn,ytn)(x^{n}_{t},y^{n}_{t}) denotes the normalized object coordinate centre and wtnw^{n}_{t} and htnh^{n}_{t} denote the normalized width and height, respectively. For the collision feature between the n1n_{1}-th object and n2n_{2}-th objet at the kk-th frame, we define it to be fn1,n2,kc=fn1,n2,ku∣∣fn1,n2,klocf^{c}_{n_{1},n_{2},k}=f^{u}_{n_{1},n_{2},k}||f_{n_{1},n_{2},k}^{loc}, where fn1,n2,kuf^{u}_{n_{1},n_{2},k} is the ResNet feature of the union region of the n1n_{1}-th and n2n_{2}-th objects at the kk-th frame and fn1,n2,klocf^{loc}_{n_{1},n_{2},k} is a spatial embedding for correlations between bounding box trajectories. We define fn1,n2,kloc=IoU(sn1,sn2)∣∣(sn1−sn2)∣∣(sn1×sn2)f^{loc}_{n_{1},n_{2},k}=\text{IoU}(s_{n_{1}},s_{n_{2}})||(s_{n_{1}}-s_{n_{2}})||(s_{n_{1}}\times s_{n_{2}}), which is the concatenation of the intersection over union (IoU), difference (−-) and multiplication (×\times) of the normalized trajectory coordinates for the n1n_{1}-th and n2n_{2}-th objects centering at the kk-th frame. We padding fn1,n2,kuf^{u}_{n1,n2_{,}k} with a zero vector if either the n1n_{1}-th or the n2n_{2}-th objects doesn’t appear at the kk-th frame.

Appendix C Dynamic Predictor

To predict the locations and RGB patches, the dynamic predictor maintains a directed graph ⟨V,D⟩=⟨{vn}n=1N,{dn1,n2}n1=1,n2=1N,N⟩\left\langle V,D\right\rangle=\left\langle\{v_{n}\}_{n=1}^{N},\{d_{n_{1},n_{2}}\}_{n_{1}=1,n_{2}=1}^{N,N}\right\rangle. The nn-th vertex ono_{n} is represented by the concatenation of its normalized coordinates btn=[xtn,ytn,wtn,htn]b^{n}_{t}=[x^{n}_{t},y^{n}_{t},w_{t}^{n},h_{t}^{n}] and RGB patches ptnp^{n}_{t}. The edge dn1,n2d_{n_{1},n_{2}} is represented by the concatenation of the normalized coordinate difference btn1−btn2b^{n_{1}}_{t}-b^{n_{2}}_{t}. To capture the object dynamics, we concatenate the features over a small history window. To predict the dynamics at the k+1k+1 frame, we first encode the vertexes and edges

where ∣∣|| indicates concatenation, ww is the history window size, fOencf_{O}^{enc} and fRencf_{R}^{enc} are CNN-based encoders for objects and relations. ww is set to 3. We then update the object influences {hn,kl}n=1N\{h_{n,k}^{l}\}_{n=1}^{N} and relation influences {en1,n2,kl}n1=1,n2=1N,N\{e_{n_{1},n_{2},k}^{l}\}_{n_{1}=1,n_{2}=1}^{N,N} through LL propagation steps. Specifically, we have

where l∈[1,L]l\in[1,L], denoting the ll-th step, fOf_{O} and fRf_{R} denote the object propagator and relation propagator, respectively. We initialize hn,to=0h_{n,t}^{o}=\textbf{0}. We finally predict the states of objects and relations at the kk+1 frame to be

where fO1predf_{O_{1}}^{pred} and fO2predf_{O_{2}}^{pred} are predictors for the normalized object coordinates and RGB patches at the next frame. We optimize this dynamic predictor by mimizing the L2\mathcal{L}_{2} distance between the predicted b^k+1n\hat{b}_{k+1}^{n}, p^k+1n\hat{p}_{k+1}^{n} and the real future locations bk+1n{b}_{k+1}^{n} and extracted patches pk+1n{p}_{k+1}^{n}.

During inference, the dynamics predictor predicts the locations and patches at kk+1 frames by using the features of the last ww observed frames in the original video. We get the predictions at the kk+2 frames by feeding the predicted results at the kk+1 frame to the encoder in Eq. 4. To get the counterfactual scenes where the nn-th object is removed, we use the first ww frames of the original video as the start point and remove the nn-th vertex and its associated edges of the input to predict counterfactual dynamics. Iteratively, we get the predicted normalized coordinates {b^k′n}n=1,k′=1N,K′\{\hat{b}_{k^{\prime}}^{n}\}_{n=1,k^{\prime}=1}^{N,K^{\prime}} and RGB patches {p^k′n}n=1,k′=1N,K\{\hat{p}_{k^{\prime}}^{n}\}_{n=1,k^{\prime}=1}^{N,K} at all predicted K′K^{\prime} frames.

Appendix D Program Parser

Following Yi et al. (2020), we use a seq2seq model (Bahdanau et al., 2015) with attention mechanism to word sequences into a set of symbolic programs and treat questions and choices, separately. The model consists of a Bi-LSTM (Graves et al., 2005) to encode the word sequences into hidden states and a decoder to attentively aggregate the important words to decode the target program. Specifically, to encode the word embeddings{wi}i=1I\{w_{i}\}_{i=1}^{I} into the hidden states, we have

where II is the number of words and fwencf_{w}^{enc} is an encoder for word embeddings. To decode the encoded vectors {ei}i=1I\{e_{i}\}_{i=1}^{I} into symbolic programs {pj}j=1J\{p_{j}\}_{j=1}^{J}, we have

where ei=ei→∣∣ei←e_{i}=\overrightarrow{e_{i}}||\overleftarrow{e_{i}} and JJ is the number of programs. The dimension of the word embedding and all the hidden states is set to 300 and 256, respectively.

Appendix E CLEVRER Operations and Program Execution

We list all the available data types and operations for CLEVRER VQA (Yi et al., 2020) in Table 8 and Table 7. In this section, we first introduce how we represent the objects, events and moments in the video. Then, we describe how we quantize the static and dynamic concepts and perform temporal and causal reasoning. Finally, we summarize the detailed implementation of all operations in Table 9.

Object and Event Concept Quantization.

We first introduce how DCL quantizes different concepts by showing an example how DCL quantizes the static object concept cube. Let fnvf_{n}^{v} denote the latent visual feature for the nn-th object in the video, SA denotes the set of all static attributes. The concept cube is represented by a semantic vector sCube{s}^{Cube} and an indication vector icubei^{cube}. icubei^{cube} is of length ∣SA∣|\texttt{SA}| and L-1 normalized, indicating concept Cube belongs to the static attribute Shape. We compute the confidence scores that an object is a Cube by

where δ\delta and λ\lambda denotes the shifting and scaling scalars and are set to 0.15 and 0.2, respectively. cos()cos() calculates the cosine similarity between two vectors and msam^{sa} denotes a linear transformation, mapping object features into the concept representation space. We get a vector of length NN by applying this concept filter to all objects, denoted as ObjFilter(cube){ObjFilter(cube)}.

Temporal and causal Reasoning.

For Filter_order of eventstype, we first filter all the valid events by find events who eventtype>ηevent^{type}>\eta. η\eta is simply set to 0 and type∈{in,out,collision}type\in\{in,out,collision\}. We then sort all the remain events based on ttypet^{type} to find the target event.

For Filter_ancestor of a collision event, we first predict valid events by finding eventstype>η\textit{events}_{type}>\eta. We then return all valid events that are in the causal graphs of the given collision event.

We summarize the implementation of all operations in Table 9.

Appendix F Trajectory Performance Evaluation.

In this section, we compare different kinds of methods for generating object trajectory proposals. Greedy+IoU denotes the method used in (Gkioxari & Malik, 2015), which adopts a greedy Viterbi algorithm to generate trajectories based on IoUs of image proposals in connective frames. Greedy+IoU+Attr. denotes the method adopts the greedy algorithm to generate trajectory proposals based on the IoUs and predicted static attributes. LSM+IoU denotes the method that we use linear sum assignment to connect the image proposals based on IoUs. LSM+IoU+Attr. denotes the method we use linear sum assignment to connect image proposals based on IoUs and predicted static attributes. LSM+IoU+Attr.+KF denotes the method that we apply additional Kalman filtering (Kalman, 1960; Bewley et al., 2016; Wojke et al., 2017) to LSM+IoU+Attr.. We evaluate the performance of different methods by compute the IoU between the generated trajectory proposals and the ground-truth trajectories. We consider it a “correct” trajectory proposal if the IoU between the proposal and the ground-truth is larger than a threshold. Specifically, two metrics are used evaluation, precision=NcorrectNpprecision=\frac{N_{\text{correct}}}{N_{p}} and recall=NcorrectNgtrecall=\frac{N_{\text{correct}}}{N_{gt}}, where NcorrectN_{\text{correct}}, NpN_{p} and NgtN_{gt} denotes the number of correct proposals, the number of proposals and the number of ground-truth objects, respectively.

Table 10 list the performance of different thresholds. We can see that Greedy+IoU achieve bad performance when the IoU threshold is high while our method based on linear sum assignment and static attributes are more robust. Empirically, we find that linear sum assignment and static attributes can help distinguish close object proposals and make the correct image proposal assignments. Similar to normal object tracking algorithms (Bewley et al., 2016; Wojke et al., 2017), we also find that adding additional Kalman filter can further slightly improve the trajectory quality.

Appendix G Statistics for CLEVRER-Grounding and CLEVRER-Retrieval

We simply use the videos from original CLEVRER training set as the training videos for CLEVRER-Grounding and CLEVRER-Retrieval and evaluate their performance on the validation set. CLEVERER-Grounding contains 10.2 expressions for each video on average. CLEVERER-Retrieval contains 7.4 expressions for each video in the training set. We for evaluating the video retrieval task on the validation set. We evaluate the performance of CLEVRER-Grounding task on all 5,000 videos from the original CLEVRER validation set. For CLEVERER-Retrieval, We additionally generate 1,129 unique expressions from the validation set as query and treat the first 1,000 videos from CLEVRER validation set as the gallery. We provide more examples for CLEVRER-Grounding and CLEVRER-Retrieval datasets in Fig. 5, Fig. 6 and Fig. 7. It can be seen from the examples that the newly proposed CLEVRER-Grounding and CLEVRER-Retrieval datasets contain delicate and compositional expressions for objects and physical events. It can evaluate models’ ability to perform compositional temporal and causal reasoning.

Appendix H Training Objectives

In this section, we provide the explicit training objectives for each module. We optimize the feature extractor and the concept embeddings in the executors by question answering. We treat each option of a multiple-choice question as an independent boolean question during training and we use different loss functions for different question types. Specifically, we use cross-entropy loss to supervise open-ended questions and use mean square error loss to supervise counting questions. Formally, for open-ended questions, we have

where CC is the size of the pre-defined answer set, pcp_{c} is the probability for the cc-th answer and yay_{a} is the ground-truth answer label. For counting questions, we have

where zz is the predicted number and yay_{a} is the ground-truth number label.

We train the program parser with program labels using cross-entropy loss,

where JJ is the size of the pre-defined program set, pjp_{j} is the probability for the jj-th program and ypy_{p} is the ground-truth program label.

We optimize the dynamic predictor with mean square error loss. Mathematically, we have

where bnb^{n} is the object coordinates for the nn-th object, pi1,i2np^{n}_{i_{1},i_{2}} is the pixel value of the nn-th object’s cropped patch at (i1,i2)(i_{1},i_{2}), and NpN_{p} is the cropped size. b^n\hat{b}^{n} and p^i1,i2n\hat{p}_{i_{1},i_{2}}^{n} are the dynamic predictor’s predictions for bnb^{n} and pi1,i2np^{n}_{i_{1},i_{2}}.