Self-Supervised Learning by Cross-Modal Audio-Video Clustering

Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, Du Tran

Introduction

Do we need to explicitly name the actions of “laughing” or “sneezing” in order to recognize them? Or can we learn to visually classify them without labels by associating characteristic sounds with these actions? Indeed, a wide literature in perceptual studies provides evidence that we rely heavily on hearing sounds to make sense of actions and dynamic events in the visual world. For example, objects moving together are perceived as bouncing off each other when the visual stimulus is accompanied by a brief sound , and the location and timing of sounds are leveraged as important cues to direct our spatiotemporal visual attention . The influence of hearing sounds in visual perception is also suggested by perceptual studies showing that individuals affected by profound deafness exhibit poorer visual perceptual performance compared to age-matched hearing controls .

In this work, we investigate the hypothesis that spatiotemporal models for action recognition can be reliably pretrained from unlabeled videos by capturing cross-modal information from audio and video. The motivation for our study stems from two fundamental challenges facing a fully-supervised line of attack to learning video models. The first challenge is the exorbitant cost of scaling up the size of manually-labeled video datasets. The recent creation of large-scale action recognition datasets has undoubtedly enabled a major leap forward in video models accuracies. However, it may be argued that additional significant gains by dataset growth would require scaling up existing labeled datasets by several orders of magnitude. The second challenge is posed by the unclear definition of suitable label spaces for action recognition. Recent video datasets differ substantially in their label spaces, which range from sports actions to verb-noun pairs for kitchen activities . This suggests that the definition of the “right” label space for action recognition, and more generally for video understanding, is still very much up for debate. It also implies that finetuning models pretrained on large-scale labeled datasets is a suboptimal proxy for learning models for small- or medium-size datasets due to the label-space gap often encountered between source and target datasets.

In this paper, we present three approaches for training video models from self-supervised audio-visual information. At a high-level, the idea behind all three frameworks is to leverage one modality (say, audio) as a supervisory signal for the other (say, video). We posit that this is a promising avenue because of the simultaneous synergy and complementarity of audio and video: correlations between these two modalities make it possible to perform prediction from one to the other, while their intrinsic differences make cross-modal prediction an enriching self-supervised task compared to within-modality learning. Specifically, we adapt the single-modality DeepCluster work of Caron et al. to our multi-modal setting. DeepCluster was introduced as a self-supervised procedure for learning image representation. It alternates between unsupervised clustering of image features and using these cluster assignments as pseudo-labels to revise the image representation. In our work, the clusters learned from one modality are used as pseudo-labels to refine the representation for the other modality. In two of our approaches—Multi-Head Deep Clustering (MDC) and Concatenation Deep Clustering (CDC)—the pseudo-labels from the second modality are supplementary, i.e., they complement the pseudo-labels generated in the first modality. The third approach—Cross-Modal Deep Clustering (XDC)—instead uses the pseudo-labels from the other modality as an exclusive supervisory signal. This means that in XDC, the audio clusters drive the learning of the video representation and vice versa. Our experiments support several interesting conclusions:

All three of our cross-modal methods yield representations that generalize better to the downstream tasks of action recognition and audio classification, compared to their within-modality counterparts.

XDC (i.e., the cross-modal deep clustering relying on the other modality as an exclusive supervisory signal) outperforms all the other approaches. This underscores the complementarity of audio and video and the benefits of learning label-spaces across modalities.

Self-supervised cross-modal learning with XDC on a large-scale video dataset yields an action recognition model that achieves higher accuracy when finetuned on HMDB51 or UCF101, compared to that produced by fully-supervised pretraining on Kinetics. To the best of our knowledge, this is the first method to demonstrate that self-supervised video representation learning outperforms large-scale fully-supervised pretraining for action recognition. Moreover, unlike previous self-supervised methods that are only pretrained on curated data (e.g., Kinetics without action labels), we also report results of XDC pretrained on a large-scale uncurated video dataset.

Related work

Early unsupervised representation learning. Pioneering works include deep belief networks , autoencoders , shift-invariant decoders , sparse coding algorithms , and stacked ISAs . While these approaches learn by reconstructing the input, our approach learns from a self-supervised pretext task by generating pseudo-labels for supervised learning from unlabeled data.

Self-supervised representation learning from images and videos. Several pretext tasks exploit image spatial context, e.g., by predicting the relative position of patches or solving jigsaw puzzles . Others include creating image classification pseudo-labels (e.g., through artificial rotations or clustering features ), colorization , inpainting , motion segmentation , and instance counting . Some works have extended image pretext tasks to video . Other video pretext tasks include frame ordering , predicting flow or colors , exploiting region correspondences across frames , future frame prediction , and tracking . Unlike this prior work, our model uses two modalities: video and audio.

Cross-modal learning and distillation. Several works train a fully-supervised encoder on one modality and distill its discriminative knowledge to an encoder of a different modality. Other works learn from unlabeled data for a specific target task . Unlike these methods, our work is purely self-supervised and aims at learning representations that transfer well to a wide range of downstream tasks. Previous cross-modal self-supervised methods most relevant to our work include audio-visual correspondence , deep aligned representations , audio-visual temporal synchronization , contrastive multiview coding , and learning image representations using ambient sound . While use only a single frame, we use a video clip. Unlike our method, clusters handcrafted audio features and does not iterate on the pseudo-labels. require constructing positive/negative examples for in-sync and out-of-sync video-audio pairs. This sampling strategy makes these approaches more difficult to scale compared to ours, as many potential out-of-sync pairs can be generated, yielding largely different results depending on the sampling choice . Recent works, such as MIL-NCE and CBT , learn from unlabeled instructional videos using text from ASR, while our approach makes use of the audio signal instead.

Technical approach

Here, we briefly discuss previous work on single-modality deep clustering in images . Then, we introduce our three multi-modal deep clustering frameworks for representation learning (Figure 1).

2 Multi-modal deep clustering

Multi-Head Deep Clustering (MDC). This model builds on SDC by adding a second classification head supervised by the other modality. Thus, in this model, each encoder has two classification heads. At each deep clustering iteration, MDC uses the cluster assignments of FvF_{v} as pseudo-labels for one head and that of FaF_{a} as pseudo-labels for the other head. Thus, each encoder needs to predict the cluster assignments of its own modality (as in SDC), but also those generated by the other modality.

Concatenation Deep Clustering (CDC). This model performs clustering of joint visual and audio features. Specifically, at each deep clustering iteration, CDC clusters vectors obtained by concatenating the visual and audio feature vectors, separately l2l_{2}-normalized. Then, it uses the resulting cluster assignments as pseudo-labels to update the weights of both EvE_{v} and EaE_{a}.

Cross-Modal Deep Clustering (XDC). Each encoder in this model relies exclusively on the clusters learned from the other modality as the supervisory signal. At each deep clustering iteration, XDC clusters the audio deep features, FaF_{a}, and uses their cluster assignments as pseudo-labels to train the visual encoder, EvE_{v}. Vice versa, XDC supervises EaE_{a} with the cluster assignments of FvF_{v}.

Experiments

Pretraining datasets. We use four datasets: Kinetics , AudioSet , IG-Kinetics , and IG-Random, which have 240240K, 22M, 6565M, and 6565M training videos, respectively. As our approach is self-supervised, thus the labels from the first three datasets are not used during pretraining. While Kinetics and AudioSet are supervised benchmarks for action recognition and audio classification, IG-Kinetics is a weakly-supervised dataset collected from a social media website using tags related to Kinetics actions. IG-Random is an uncurated dataset of random videos from the same website. Videos are 1010-second long in Kinetics and AudioSet and 1010-to-6060-second long in IG-Kinetics and IG-Random. We filter out around 77K Kinetics videos that have no audio. Furthermore, we randomly sample 240240K videos from AudioSet and denote this subset as AudioSet-240240K. We generate this subset to have AudioSet data of the same size as Kinetics, in order to study the effects of pretraining with the same data size but on a different data distribution and domain.

Downstream datasets. We evaluate our pretraining performance on three downstream benchmarks: UCF101 , HMBD51 , and ESC50 , which have 1313K, 77K, and 22K examples from 101101, 5151, and 5050 classes, respectively. UCF101 and HMDB51 are action recognition datasets, while ESC50 is a sound classification dataset. UCF101 and HMDB51 have 33 official train/test splits, while ESC50 has 55 splits. We conduct our ablation study (Subsection 4.2) using split-1 of each dataset. We also report our average performance over all splits when we compare with state-of-the-art methods in Section 6.

Baselines. We consider two baselines: Scratch and Supervised Pretraining (Superv). The first is a randomly-initialized model trained from scratch directly on the downstream task, while the second is a model pretrained in a supervised fashion on a large labeled dataset (e.g., Kinetics) and then finetuned on the downstream task. We note that these two baselines are commonly regarded as the lower and upper bounds to gauge the quality of self-supervised representation learning methods .

Backbone architectures. We employ R(2+1)D and ResNet as EvE_{v} and EaE_{a}, respectively. EvE_{v}’s input is a 3<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>L<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>H<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>W3<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>L<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>H<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>W clip, where 33 refers to the RGB channels, LL is the number of frames, and HH and WW are the frame height and width. EaE_{a}’s input is a Q<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>PQ<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>P spectrogram image extracted from the audio signal, where QQ is the number of MEL filters and PP is the number of audio frames.

Pretraining and evaluation details. We choose the 1818-layer variants of R(2+1)D and ResNet encoders. We use clips of LL=88 frames for pretraining and finetuning our visual encoder EvE_{v}. We scale frames such that the smallest dimension is 256256 pixels and then random crop images of size 224<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>224224<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>224. We extract video clips at 3030 fps and employ temporal jittering during training. For the audio input, we sample 22 seconds and use QQ=4040 MEL filters and PP=100100 audio frames. For inference on the downstream tasks, we uniformly sample 1010 clips per testing example and average their predictions to make a video-level prediction. We use only one crop per clip: the center 8<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>224<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>2248<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>224<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>224 crop for video and the full 40<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>10040<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>100 crop for audio. We provide more details in the supplementary material.

2 Ablation study

Study 1: Single-modality vs. multi-modal deep clustering. This experiment compares the four models presented in Section 3. We pretrain SDC, MDC, CDC, and XDC on Kinetics and report their performance on the downstream tasks in Table 1. To better understand XDC, we also conduct a new set of baselines, called same-modality-XDC, where XDC is trained with two encoders defined on the same modality (either visual or audio). Note that all models in this ablation study use the same visual and audio encoders and only differ in the way they use self-supervision. It takes on average 55 to 66 deep clustering iterations for these models to converge. Observations: (I) The four self-supervised deep clustering models outperform the Scratch baseline on all downstream benchmarks. This shows that our self-supervised pretraining is effective and generalizes well to multiple tasks. (II) All multi-modal models (MDC, CDC, and XDC) significantly outperform SDC by up to 12.412.4%, 7.67.6%, and 11.511.5% on UCF101, HMDB51, and ESC50, respectively. This validates the importance of multi-modal modeling compared to single-modality. (III) XDC achieves the best performance across all tasks. What distinguishes XDC from the other models is that each modality encoder in XDC is self-supervised purely by the signal from the other modality. The encoders in CDC, MDC, and SDC all employ a self-supervision signal coming from the same modality. Thus, this suggests that encoders learn better when purely supervised by a different modality. We provide the following intuition on why XDC is better than CDC and MDC. XDC groups samples together when they are similar in one of the two modalities (video to supervise the audio encoder, audio to supervise the visual encoder). Instead, CDC groups samples together only if they are similar according to both the audio and the video modality (to supervise both encoders). Thus, XDC visual and audio clusters allow for more diversity than those of CDC. We hypothesize that this diversity allows XDC to learn richer representations, which translates into better performance on the downstream tasks. Also, recent work has shown that models trained on different modalities learn and generalize at different speeds, and that training them jointly (as done in MDC which uses two-modality heads) is sub-optimal. We believe that this could contribute to MDC performing worse than XDC, which optimizes for each modality independently. (IV) The same-modality-XDC baselines perform similarly to SDC and are 88-1212% worse than multi-modal-XDC. This suggests that cross-modality provides a superior supervisory signal for self-supervised learning and that multi-modal-XDC is the best model not because of its optimization strategy but rather because of the use of the other modality for pseudo-labeling. Given the results of this study, we opt to use only XDC in the rest of the experiments. Finally, to show that XDC works for different backbones, we re-do Study 1 with ResNet3D in the supplementary material.

Study 2: The number of kk-means clusters. This study explores the effects of changing the hyperparameter kk in kk-means clustering. We pretrain XDC on three datasets, Kinetics, AudioSet-240240K, and AudioSet, using kk=64,128,256,512,64,128,256,512, and 10241024 clusters (Table 4.2). Observations: (I) The best kk value is not sensitive to the number of semantic labels in the downstream datasets. For example, HMDB51 and ESC50 have about the same number of labels but different best kk value. (II) Similarly, the best kk value seems uncorrelated with the number of original semantic labels of the pretraining dataset, e.g. 400400 in Kinetics. We reiterate here that our approach is self-supervised and does not use the labels of the pretraining dataset. (III) The best kk value tends to get larger as the pretraining data size increases. For example, the best kk for HMDB51 shifts from 128128 to 256256 when moving from pretraining on AudioSet-240240K to the full AudioSet. We hypothesize that there is a more diverse sample set to cluster when the pretraining data size increases. Thus, we can have more fine-grained clusters (higher kk) and make our self-supervised classification problem harder. This aligns with previous self-supervised works that showed benefits from making the pretext task harder.

Study 3: Pretraining data type and size. Here, we investigate the effects of two pretraining characteristics: data size and type. To this end, we pretrain XDC on Kinetics (240240K examples), AudioSet-240240K (240240K examples), AudioSet (22M examples), IG-Kinetics (6565M examples), and IG-Random (6565M examples). Kinetics and IG-Kinetics videos are collected originally for activity recognition, while AudioSet videos are aimed for audio event classification. IG-Random is an uncurated/unsupervised dataset. In addition to video datasets, we also experiment with ImageNet to understand how much action recognition benefits from supervised pretraining on object classification. For ImageNet, we inflate the images into static video clips (repeating the same frame) and pretrain our video model on this dataset. Table 3 presents the results of this study. Observations: (I) XDC improves across all three downstream tasks as the pretraining data size increases. For example, XDC on HMDB51 improves by 9.8%9.8\%, 22.2%22.2\%, and 24.1%24.1\% when pretrained on AudioSet, IG-Random, and IG-Kinetics, respectively, compared to the results when pretrained on Kinetics. (II) XDC outperforms Kinetics fully-supervised pretraining by 5.1%5.1\% on HMDB51 and by 0.6%0.6\% on UCF101. To the best of our knowledge, XDC is the first method to demonstrate that self-supervision can outperform large-scale full-supervision in representation learning for action recognition. (III) The performance of the fully-supervised pretrained model is influenced by the taxonomy of the pretraining data more than the size. For example, supervised-pretraining on Kinetics gives better performance on both UCF101 and HMDB51 compared to supervised-pretraining on AudioSet (which is 88 times larger than Kinetics) and ImageNet. One the other hand, XDC performance is less sensitive to the data type, as it implicitly learns the label space rather than depend on a space manually defined by annotators.

Study 4: Curated vs. uncurated pretraining data. The overarching goal of self-supervised representation learning is to learn from the massive amounts of unlabeled data. Previous self-supervised methods have pretrained on videos from supervised (curated) datasets (e.g., Kinetics) without using the labels. However, even without using labels, those videos are still biased due to the sampling distribution (e.g., taxonomy of the curated dataset). To this end, we study the effects of self-supervised representation learning from uncurated data. Table 4.2 compares XDC pretrained on IG-Kinetics (curated, as videos were tag-retrieved) vs. IG-Random (uncurated) using 11M, 1616M, and 6565M videos. Observations: (I) Curated pretraining gives better results on UCF101 and HMDB51, while uncurated pretraining is better on ESC50 at large scale. We hypothesize that the bias of IG-Kinetics towards semantics of human actions is the reason behind the positive effect of curation on UCF101 and HMDB51. However, such bias negatively impacts the performance on ESC50. (II) The performance gap between the curated and uncurated pretraining shrinks significantly as we increase the data size. For example, the performance gap on HMDB51 drops from 5.2%5.2\% to 2.1%2.1\% and 1.9%1.9\% when the pretraining size increases from 11M to 1616M and 6565M videos, respectively. This implies that XDC can learn meaningful representations from truly uncurated data. To the best of our knowledge, XDC is the first self-supervised method to study pretraining on large-scale uncurated video data.

Study 5: Full finetuning vs. learning fc-only. Here, we study two approaches for transferring XDC to downstream tasks. Full finetuning: we finetune all parameters of the pretrained encoder on the downstream task. Learning fc-only: we fix the pretrained encoder and learn a linear classifier for the downstream task, i.e., a fully-connected (fc) layer on top of the frozen features. Table 5 compares XDC with the supervised pretrained approaches under these two transfer-learning schemes. Observations: (I) The accuracy of most pretrained models (fully-supervised or self-supervised) degrades, when used as a fixed feature extractor compared to when they are fully-finetuned on the downstream tasks. Nonetheless, the relative performance of XDC compared to supervised pretrained models stays generally the same when fully vs. fc-only finetuned on the downstream task. This suggests that XDC pretraining is useful both as a fixed feature extractor and as a pretraining initialization. (II) XDC as a fixed feature extractor outperforms many fully-finetuned supervised models. For example, fc-only XDC outperforms, by significant margins, the fully-finetuned supervised AudioSet- and ImageNet-pretrained models on both UCF101 and HMDB51. (III) We observe that fully-supervised pretraining, followed by fc-only finetuning performs well when the pretraining taxonomy is well aligned with that of the downstream task. For example, pretraining on Kinetics by learning fc-only on HMDB51 and UCF101 gives the best performance. This is expected as the label spaces of HMBD51 and UCF101 overlap largely with that of Kinetics. This suggests that fully-supervised pretraining is more taxonomy/downstream-task dependent, while our self-supervised XDC is taxonomy-independent.

What does XDC actually learn? What semantic signals does the algorithm use to train its encoders? Here, we try to answer these questions by inspecting the kk-means clustering results produced by the last iteration of XDC. Figure 2 visualizes some audio and video clusters learned by XDC when trained on Kinetics. These clusters are the top 22 audio clusters (left) and the top 22 video clusters (right) ranked by purity w.r.t. Kinetics action labels. More clusters are presented in Table 6. We observe that the top-purity clusters learned from both modalities exhibit strong semantic coherence. For example, the audio 1st and 8th ranked clusters include concepts related to playing musical instruments that have similar sounds, while the 1st ranked video cluster also groups playing-instrument concepts, but mainly because of their appearance, as the cluster is all about guitars. Other interesting clusters include: grouping by motor-engine sounds (audio #10), by different swimming strokes (video #4), by different golf shots (video #5), and different cooking activities (video #10). In the bottom-ranked clusters, although the purity w.r.t. Kinetics concepts is low, we still find some coherence, mostly at the scene level: a farm setting in audio #127 (“grooming horse”, “milking cow”) and gym activities in video #63 (“pull ups”, “punching bag”). Many other bottom-ranked clusters appear to lack semantic coherence when viewed through the lens of Kinetics labels. However, one of the motivations behind the design of self-supervised methods is precisely to bypass the hand-design of label spaces, which may not be the optimal ones for general representation learning. Our experiments suggest that the label space learned by XDC yields strong and general audio and video features even though it does not align perfectly with the taxonomies of existing datasets.

Experimental setup. Here, training is similar to our ablations except that we re-train our video encoder on the last clustering assignment using 3232-frame clips. Then following , we finetune on UCF101 and HMDB51 using 3232-frame clips for both XDC and the fully-supervised baselines. Inference is similar to our ablations except for using 3232-frame clips. For the audio event classification dataset DCASE , we follow and extract conv_\_5 features for 6060 uniformly-sampled clips per audio sample and learn a linear SVM. We report the average top-1 accuracy over all splits.

Video action recognition. Table 7(a) compares XDC pretrained on four large-scale datasets against state-of-the-art self-supervised methods, after finetuning on the UCF101 and HMDB51 benchmarksAll XDC pretrained models are publicly released on our project website.. We also compare against two fully-supervised methods pretrained on ImageNet and Kinetics. Results: (I) XDC pretrained on IG-Kinetics sets new state-of-the-art performance for self-supervised methods on both benchmarks, outperforming Elo by 1.7%1.7\% on UCF101 and 1.5%1.5\% on HMDB51. Moreover, XDC significantly outperforms fully-supervised pretraining on Kinetics: by 1.3%1.3\% on UCF101 and by 3.8%3.8\% on HMDB51. (II) When directly compared on the same R(2+1)D-18 architecture, XDC pretrained on Kinetics slightly outperforms AVTS by 0.6%0.6\% on UCF101 and 0.3%0.3\% on HMDB51. However, when both methods are pretrained on AudioSet, XDC outperforms AVTS with larger margins: by 3.9%3.9\% on UCF101 and 5.6%5.6\% on HMDB51. This shows that XDC scales better than AVTS. To further verify that XDC scales better, we pretrained AVTS on AudioSet-240240K using R(2+1)D-18 and got 76.9%76.9\% and 40.7%40.7\% for UCF101 and HMDB51 on split-1, showing a smaller margin between XDC and AVTS than when both are pretrained on the full AudioSet (cf. Table 3).

Audio event classification. Table 7(b) compares XDC pretrained on AudioSet and IG-Random against the state-of-the-art self-supervised methods for audio classification. XDC achieves state-of-the-art performance on DCASE and competitive results on ESC50 with only a 1.1%1.1\% gap with .

In this section, we further demonstrate that XDC can be useful beyond video and audio classification. In particular, we employ the recent G-TAD action localization algorithm, where we replace the clip features (originally extracted from a TSN model pretrained on Kinetics) with our XDC features from the R(2+1)D-18 model pretrained on IG-Kinetics or IG-Random. We compare against the features from the R(2+1)D-18 model fully-supervised pretrained on Kinetics. We emphasize that we do not finetune any of the feature extractors used in this experiment. We follow the default hyperparameters setting of G-TAD. Table 8 shows temporal action localization results of G-TAD with different features on THUMOS14 dataset. It reports the mean Average Precision (mAP) results at different temporal Intersection over Union (tIoU) thresholds. Both XDC variants outperform the fully-supervised features across all tIoU thresholds. This confirms the same trend observed in tasks presented in Section 6 and suggests that XDC can be used for other tasks.

We presented Cross-Modal Deep Clustering (XDC), a novel self-supervised model for video and audio. XDC outperforms not only existing self-supervised methods but also fully-supervised ImageNet- and Kinetics-pretraining for action recognition. To the best of our knowledge, XDC is the first to show self-supervision outperforming large-scale full-supervision pretraining for action recognition when pretrained on the same architecture and a larger number of uncurated videos.

Video has become a commonplace in society. Its uses range from entertainment, to communication and teaching. Thus, the learning of semantic representations of video has broad and far-reaching potential applications. The authors do not foresee major ethical issues associated to this work. However, as the proposed approach is self-supervised, it will learn the inherent properties and structure of the training data. Thus, the learned model may exhibit biases intrinsically present in the data.

The authors thank Mengmeng Xu for his valuable help with the THUMOS14 experiments. The authors appreciate the anonymous NeurIPS reviewers for their constructive feedback. Humam Alwassel was partially supported by the King Abdullah University of Science and Technology (KAUST) Office of Sponsored Research (OSR) under Award No. OSR-CRG2017-3405.

Supplementary Material

Appendix A Optimization challenges

In this section, we give the details of the full optimization cycle and discuss differences between the single-modality baseline and our multi-modal models.

Trivial solutions. As discussed in , SDC may converge to trivial solutions, corresponding to empty clusters or encoder parameterizations, where the classifier predicts the same label regardless of the input. DeepCluster proposes workarounds to tackle these issues, involving reassigning empty cluster centers and sampling training images uniformly over the cluster assignments. While these strategies mitigate the issues, they do not fix the main cause of the problem: SDC learns a discriminative classifier on the same input from which it learns the labels. On the other hand, our multi-modal deep clustering models are less prone to trivial solutions because they learn the discriminative classifier on one modality and obtain the labels from a different modality. In our training, we never encountered the issue of empty clusters or few-class predictions for any of our multi-modal clustering approaches.

Initialization and convergence. Our initial pseudo-labels come from clustering features of randomly-initialized encoders. Such pseudo-labels are “good enough” to capture some weak similarities between the input samples as features from randomly-weighted networks have shown decent results on image and audio classification . Another potential option involves generating the initial pseudo-labels by clustering hand-crafted features, e.g. iDT and audio spectrograms. Hand-crafted features capture low-level semantics that may help the encoders learn better or faster. Indeed, in small-scale experiments, we observed that clustering handcrafted features in the initial iteration reduces the number of clustering iterations needed to learn a well-performing encoder. However, we decided to not pursue this further, since these features are computationally expensive to extract and thus are not suitable for large-scale training on millions of examples. Furthermore, handcrafted features may bias the learning to reflect the design choices behind these manually-engineered descriptors.

Clustering and optimization schedule. Following previous work , we cluster the deep features using the kk-means algorithm primarily for its desirable properties of efficiency and scalability. The number of kk-means clusters is a key hyperparameter in our framework. Intuitively, using more clusters makes the pretext task harder, as it increases the number of pseudo-classes the classifier must recognize. On the other hand, the diversity of samples to cluster effectively dictates the maximum kk, for which the grouping is still sensible. Taking into account these factors, we explore the effects of kk in our ablation study in Subsection 4.2 of the main manuscript. Another important hyperparameter of our framework is the number of training epochs for the encoders, before re-clustering the learned features. DeepCluster re-clusters after each epoch, which is an expensive design choice when scaling to millions of training samples. Thus, we choose to fix the pseudo-labels and train the encoders until the validation loss for predicting the pseudo-labels saturates. Then, we re-cluster the newly learned features, reassign pseudo-labels, reset the classification layer, and repeat the same process. We find this strategy to be more efficient, as it reduces the number of times we need to invoke kk-means.

Appendix B Learning using audio rather than text from ASR

We note that while our approach was demonstrated by leveraging audio, the method is general and is easy to adapt to other modalities, including text. While video and text are semantically correlated, audio and video are temporally correlated. Thus, these two form of correlations are likely to provide different forms of self-supervision, potentially leading to further gains when used in combination. A disadvantage of text from ASR is that it is only available for videos with speech. Audio provides information about environmental sounds beyond speech (e.g. walking steps, playing guitar, and dog barking) and allows us to train on uncurated datasets of arbitrary Web videos, as we demonstrated with IG-Random.

Appendix C Hyperparameters and training details

Training. We train our models using caffe2 with distributed SGD on a GPU cluster, and employ the warmup scheme proposed in . The main training parameters are presented in Table 9. We note that the epoch size can be different from the actual number of videos. This is because the total number of clips the model sees during training (with temporal jittering) can be larger than the number of videos.

Pretraining parameters. We pretrain XDC and other baselines using the parameters described in Table 10. Early stopping is used for pretraining on small datasets such as Kinetics and AudioSet to stop before the model starts overfitting on the pretext task. For IG-Kinetics and IG-Random, we do not observe overfitting. We pretrain XDC on IG-Kinetics and IG-Random longer in the last deep clustering iteration (denoted as IG-Kinetics* and IG-Random* in Table 10). When pretraining our R(2+1)D on longer clips (e.g. 32 frames), due to the GPU memory limit, we reduce the mini-batch size to 88 (instead of 3232) and the base learning rate to 0.00250.0025 (instead of 0.010.01).

Finetuning parameters. We provide finetuning hyperparameters in Table 11. Different pretraining methods may have different optimal base learning rate when finetuned on downstream tasks. Thus to make a fair comparison, we cross-validate the finetuning using the same set of base learning rates (presented in Table 12) and report the best result for each pretraining method. As we observed that higher learning rates tend to be beneficial when learning FC-only, we use a wider set of learning rates to cross-validate FC-only models. As done during pretraining, when finetuning R(2+1)D on longer clips (i.e. 32 frames), we reduce the mini-batch size to 88 and reduce the base learning rate to 1/41/4 of its original rate.

Appendix D XDC using a different backbone architecture

We pretrain XDC on Kinetics with ResNet3D-18 as the visual backbone and keep the same audio encoder (ResNet-18). The results are compared with those of baselines in Table 13. XDC with the ResNet3D-18 backbone outperforms the training from scratch baseline by good margins on three downstream tasks.

Appendix E Additional qualitative results

XDC clusters. Tables 14 and 15 present the top and bottom 10 audio and video clusters learned with XDC on Kinetics, ranked by their purity with respect to Kinetics labels. We list the 55 most frequent concepts of each cluster.

XDC filters. Figure 3 visualizes and compares conv_1 spatial and temporal filters of R(2+1)D learned by self-supervised XDC pretraining on IG-Kinetics versus fully-supervised pretraining on Kinetics. We observe some differences in both spatial and temporal filters between XDC and fully-supervised pretraining. In particular, XDC learns a more diverse set of motion filters.