Sound Localization by Self-Supervised Time Delay Estimation

Ziyang Chen, David F. Fouhey, Andrew Owens

Introduction

Sounds in the world arrive at one of our two ears slightly sooner than the other. This interaural time delay, which generally lasts only a few hundred microseconds, indicates a sound’s direction and thus provides an important cue for multimodal perception. In humans, for example, time delays convey the positions of objects that move out of sight, and are integrated with visual cues when localizing events . Visual information can also guide the sound localization process, allowing us to find a particular event of interest through binaural cues, while ignoring the others.

While high-quality stereo sound recordings are now abundant, such as in the audio tracks of videos recorded by consumer phones, existing methods often struggle to localize sound sources within them, particularly when they contain correlated noise or multiple sound sources. The localization problem has typically been addressed by matching hand-crafted features and, recently, by supervised learning . However, the difficulty in acquiring natural labeled data has limited their effectiveness. Many approaches, consequently, resort to using simulated training data that may not be fully representative of the world.

We propose to address these problems by learning time delay estimation from real, unlabeled recordings. We take inspiration from work in self-supervised visual tracking that learns space-time correspondences from videos, such as through cycle consistency and contrastive learning . Analogously, our approach is based on learning audio embeddings that can be used to find interaural correspondences: pairs of sounds from different stereo channels that correspond to the same underlying events.

We introduce a model, inspired by the contrastive random walk of Jabri et al. , that learns cycle consistent features from unlabeled stereo sound. This model maximizes the return probability of a random walk on a graph whose nodes correspond to the audio samples in each channel. In this graph, edges connect samples between channels, and the walk’s transition probabilities are defined by learned embeddings. We show examples of time delay estimates for two real-world videos in Figure 1.

We also propose a model inspired by instance discrimination that can perform a novel visually-guided time delay estimation task: localizing a speaker in a multi-speaker audio recording, given only their visual appearance. The resulting model is simple and can accurately localize speakers, without the need for explicitly separating sounds in the mixture. We also use this approach to train audio-based localization models solely from mono audio, which in some domains may be more readily available than stereo sound. This model uses data augmentation to incorporate knowledge about invariances to important sources of variation.

Through experiments on simulated environments with metrically accurate ground truth, and on internet videos with directional judgments annotated by human listeners, we show:

Interaural time delays can be accurately estimated through self-supervised learning, using either unlabeled stereo and mono training data.

Our models provide robustness to distracting sounds within a mixture, and perform well on real-world recordings, obtaining competitive performance with state-of-the-art supervised methods.

Visual signals allow our models to localize specific speakers within mixtures.

Related Work

Humans use two main cues for estimating the azimuth of a sound: interaural time differences (ITD) and interaural intensity differences (IID), i.e., the difference in the loudness of the sounds entering both ears . In practice, IID is primarily useful for high-frequency sounds that are close to the observer, while ITD is useful for low-frequency sounds and is relatively unaffected by distance . Our work is thus complementary to methods that use IID cues. Humans can accurately estimate azimuth from multi-source mixtures , and integrate vision with binaural cues , motivating our work on multi-speaker time delay estimation.

Time delay estimation.

Time delay estimation is a classic signal processing problem. In early work, Carter et al. estimated time delays using generalized cross-correlation with phase transform (GCC-PHAT), which corresponds to a maximum likelihood estimate under low noise . Other work uses beamforming or subspace methods . Comanducci et al. trained a convolutional network (CNN) to denoise GCC-PHAT features. Other work trains a multi-layer perceptron to predict a time delay from a raw waveform , trains recurrent networks on hand-crafted features , and uses 3D CNNs . Christensen et al. used GCC-PHAT and echolocation to estimate depth maps from audio. In concurrent work, Chen et al. localized multiple sounds by jointly solving source separation and time delay estimation problems. Time delay estimation also has a wide range of applications and modalities, such as oceanography , wireless networking , sonar , and possibly directional olfaction . In contrast, we pose time delay estimation as a self-supervised learning problem, and we do not require hand-crafted features or labels.

Supervised binaural localization.

Vecchiotti et al. estimated sound direction directly from raw waveforms. Other work uses Short-time Fourier Transform or beamforming features . Due to the challenge in obtaining labeled data, these methods have largely been trained on synthetic or lab-collected data. In contrast to these approaches, we learn a specific (but widely useful) cue—the time delay—through self-supervision on natural data.

Audio-visual binaural learning.

Yang et al. distinguished between audio-visual examples in which the stereo channels have (or have not been) swapped, resulting in a representation that can be finetuned to solve localization tasks. In contrast, our model can optionally be trained and deployed solely with audio, and produces an output—the time delay—that is directly correlated with sound direction, without the need for finetuning. Gan et al. used a car detector to provide pseudo ground truth for a sound-based localization method. Since the training data comes from a supervised car detector, the model relies on labeled training data, whereas ours is self-supervised. Later work extends this approach by distilling supervision from multiple visual classifiers and modalities. Other work generates stereo sound from mono audio using images , largely by adjusting the relative volume of the channels to simulate IID cues.

Audio-visual sound localization and separation.

A variety of methods have been proposed for using vision to localize and separate sounds. Classic work searches for cross-modal similarity in statistical models . Later work uses contrastive learning to find image regions that are highly correlated with sound , and separates sounds from synthetic mixtures . Recent work has applied the contrastive random walk to localize multiple sounds within images . This method learns correspondences between image patches and (mono) audio, whereas our model learns correspondences between the signals in each stereo channel.

Audio self-supervision.

A variety of methods have been proposed for learning audio representations through self-supervision, typically for semantic recognition tasks, such as music or speech understanding. These include contrastive learning , autoencoding , multi-task learning with pretext tasks , and generative autoregressive models . In contrast, we learn a representation for learning interaural correspondence in binaural audio.

Learning visual correspondences.

We take inspiration from methods that learn space-time correspondences from video. These include methods that colorize grayscale video , cycle-consistent feature representations and slow features . Other work has shown that features learned through instance discrimination are effective for tracking. However, these methods have not been applied to learning stereo audio correspondences. In our models, by contrast, vision is used to aid the audio matching process. We adapt several of these methods to learn correspondences between temporal samples of audio for binaural matching.

Method

where hi=h(xi){\mathbf{h}}_{i}=h({\mathbf{x}}_{i}) are the features for xi{\mathbf{x}}_{i}, and hi(t){\mathbf{h}}_{i}(t) is the dd-dimensional feature embedding for time tt.

Traditionally, the audio features, hh, are defined using hand-crafted features. For example, the widely-used Generalized Cross Correlation with Phase Transform (GCC-PHAT) whitens the audio by dividing by the magnitude of the cross-power spectral density. This approach provides the maximum likelihood solution under certain ideal, low-noise conditions .

We propose, instead, to learn hh through self-supervision from unlabeled data. These features ought to capture interaural correspondences: observations in both waveforms that were generated by the same underlying events should be close in embedding space. We consider models that can be trained solely from unlabeled stereo or mono sound (Sec. 3.1), or that learn to perform visually-guided estimation from audio-visual data (Sec. 3.2).

We propose models that learn interaural correspondence from unlabeled data.

Our embeddings should provide cycle consistent matches: the process of matching features from x1{\mathbf{x}}_{1} to those in x2{\mathbf{x}}_{2} should yield the same correspondences as matching in the opposite direction, from x2{\mathbf{x}}_{2} to x1{\mathbf{x}}_{1}. We use this idea to learn a representation from unlabeled stereo sounds.

We adapt the contrastive random walk model of Jabri et al. to binaural audio (Fig. 2a). We create a graph that contains nodes for each of the temporal sample xi(t){\mathbf{x}}_{i}(t) from both channels, with edges connecting the nodes that come from different channels.Following visual tracking work , one could potentially extend this approach to microphone arrays with 3 or more channels by performing the walk over all channels. We then perform a random walk that transitions from nodes in x1{\mathbf{x}}_{1} to those in x2{\mathbf{x}}_{2}, then back to x1{\mathbf{x}}_{1}, with transition probabilities that are defined by dot products between embedding vectors:

where Aij(s,t)A_{ij}(s,t) is the probability of transitioning from sample ss in xi{\mathbf{x}}_{i} to sample tt in xj{\mathbf{x}}_{j}, and a temperature constant cc. The features hi=h(xi;θ){\mathbf{h}}_{i}=h({\mathbf{x}}_{i};\theta) are parameterized with network weights θ\theta and are represented using a CNN (Sec. 3.3). We maximize the log return probability of a walk that moves between nodes in the two channels:

where the log⁡\log is computed element-wise. We also found it helpful to incorporate knowledge about invariances to important sources of variation, such as to noise. To do this, we also apply data augmentation to two audio channels during the walk similar to Hu et al. (see Sec. A.5 for details).

Slow features.

We also train a variation of the model that learns to associate embeddings that temporally co-occur, taking inspiration from methods that learn slow features in video and audio-visual synchronization . These pairs of embeddings are more likely (than misaligned timestamps) to correspond to the same events. We minimize:

Instance discrimination.

We also consider models that can be trained solely with mono audio using instance discrimination . In lieu of a second audio channel, we create synthetic views of mono audio, using data augmentation that encourages invariances that are likely to be useful for interaural matching. We minimize:

over all timesteps tt, where h=h(x){\mathbf{h}}=h({\mathbf{x}}) are the features for a mono audio x{\mathbf{x}}, and h^=h(x^){{{\hat{\mathbf{h}}}}}=h({{\hat{\mathbf{x}}}}) are features computed from an augmented version of x{\mathbf{x}}.

Unless otherwise specified, we perform two types of augmentation: time shifting and volume adjustment. To model the challenges in time delay estimation, we choose negative examples exclusively from x{\mathbf{x}}, rather than other examples in the batch . We ensure that augmented positive views are always taken from the corresponding timestep, i.e., we undo any time-shifting augmentation when indexing h^(t){{\hat{\mathbf{h}}}}(t).

2 Visually-guided time delay estimation

We also apply our model to the novel problem of estimating the time delay for a single sound within a mixture using visual information. Given a sound mixture containing multiple simultaneous speakers, we estimate the time delay for one object, given a visual representation of its appearance (e.g., localizing a speaker using a visual representing their face). The visual input need not co-occur with the audio. For example, the object may be off-screen, or its visual features may have been extracted at an earlier time.

over all timesteps tt, where gi=g(xi,Iu){\mathbf{g}}_{i}=g({\mathbf{x}}_{i},I_{u}) are the learned audio-visual features for channel xi{\mathbf{x}}_{i}. Here, gg obtains its embedding by fusing audio from one channel with the input image. As in the instance discrimination model, we apply augmentation to g2{\mathbf{g}}_{2}. Note that this task cannot be solved without IuI_{u}: from audio alone, the model would be unable to determine whether the true delay is τu\tau_{u} or τv\tau_{v}.

3 Learning a time delay estimation model

We now describe how these self-supervised learning models can be trained, and how they can be used to estimate time delays.

For our audio-visual model, we represent the visual information using (pretrained) FaceNet . This allows our model to estimate attributes of speakers from face crops, similar to the work in source separation that uses face embeddings . We fuse the audio and visual features after the second convolution block of the audio subnetwork by concatenating the 128-dimensional visual features at each time-frequency position.

Datasets.

We train our audio models on datasets of stereo sound: FAIR-Play , which has 1,871 videos (5.2 hours) of lab-collected music performances from a small number of rooms and Free-Music-Archive (FMA) , a dataset of 101K (841 hours) music recordings created by a large number of artists.

For the visually-guided model, we train our model on VoxCeleb2 with 500 randomly selected identities. We randomly create training mixtures from mono sound, without the speaker identity labels and without using a simulator.

Training.

We use the AdamW optimizer with a learning rate =10−4=10^{-4}, a cosine decay learning rate scheduler, a batch size of 48, a temperature c=0.05c=0.05 following , and early stopping. Please see Sec. A.5 for more training details.

Self-supervised learning formulation.

For all models, we extract our examples from a 1220-sample waveform, sampled at 16 Khz. We obtain our embeddings using a sliding window of size 0.064s (1024 samples), with a step of 4 samples, yielding 49 audio clips. We apply random stereo channel swapping and channel-wise waveform rescaling for augmentation in all models. In some experiments, we also add random noise, add reverberation, and mix in other sounds as additional augmentation. At test time, we obtain a denser audio graph by using the step of 1 sample for the sliding window. Our model can use input sounds with a variety of durations without retraining, due to the fully-convolutional network architecture . For the audio-visual task, we use a window length of 0.96s or 2.55s to obtain more temporal context, since this problem involves jointly solving a separation task.

Estimating delays from features.

After learning our representation hh, we can use it to estimate the time delay, such as by maximizing Rx1,x2R_{{\mathbf{x}}_{1},{\mathbf{x}}_{2}} (Eq. 1). We have found that this procedure can affect the quality of the prediction for both learned and hand-crafted methods (e.g., due to outliers), so we evaluate a number of different variations in our experiments. In our approach, each embedding votes on a value for τ\tau. We then choose a single time delay for the audio from these votes, either by taking the mean or by using a RANSAC-like mode estimation method. In the latter, we first select the delay with the most votes, then average the inliers (those within a small threshold of the chosen value). This vote can be performed by the nearest neighbor search, or by treating the learned similarities as probabilities (Eq. 2) and taking the expectation, i.e., 1n∑s,ττA12(s,τ)\frac{1}{n}\sum_{s,\tau}\tau A_{12}(s,\tau) .

Experiments

We evaluate our methods using both simulated audio, where time delays can be measured exactly, and real-world binaural audio from unknown microphone geometry, where quantized sound direction categories are labeled by humans.

Before considering real-world audio, we evaluate each model’s performance on time delay estimation task using simulated environments, following . While the resulting sounds are considerably simpler than real-world recordings, they allow us to obtain metrically accurate ground-truth time delays, and to systematically vary different experimental conditions, such as the amount of background noise.

Following previous work , we simulate stereo sounds using Pyroomacoustics . We create three simulated environments with rooms of different sizes and microphone positions. For our sound sources, we take speech sounds from TIMIT (recorded in anechoic conditions) and place them at random angles sampled uniformly from (-90∘, 90∘) and distances (0.5m, 3.0m) with respect to the microphone. We add independent Gaussian noise to create conditions with different signal-noise ratio (SNR) levels, and consider a variety of reverberation times (RT60\text{RT}_{60}). The ground-truth time delay can straightforwardly be calculated from the sound source and the microphone pose. This simulated test set, which we call TDE-Simulation, contains approximately 6K audio samples total (please see Sec. A.4 for more details).

Models.

We evaluated our audio-based learning methods: 1) StereoCRW, contrastive random walks trained on stereo sounds, 2) ZeroNCE, slow features trained on stereo sounds (named after VINCE ), and 3) MonoCLR, and instance discrimination trained on mono sounds (named after SimCLR ).

We compared our methods with the widely-used GCC-PHAT , a hand-crafted audio feature. We also compared with the recent supervised method Salvati et al. , which trains a CNN on parameterized GCC-PHAT features to regress time delay. We trained this model on simulated stereo sounds, based on audio clips from VoxCeleb2 to obtain human speech signals. To improve this baseline’s performance, we make a modification: in addition to the noise and reverberation augmentations from , we train with synthetic sound mixtures, in which a background sound is added to the input waveform. We regard this supervised method as an approximate upper bound for the simulation-based experiments. We provide all methods with the same duration audio as input, and evaluate different post-processing methods.

As shown in Tab. 4.1, the StereoCRW model substantially outperforms GCC-PHAT when it is trained on a large stereo dataset, FreeMusic-Archive, obtaining performance comparable with supervised models trained on synthetic data. While ZeroNCE is trained with real stereo sound, its loss implicitly assumes the true time delay is zero, which is violated in real scenes. The cycle consistency loss in StereoCRW does not make this assumption, allowing it to learn from more complex data and learn better representations. The ZeroNCE model outperforms MonoCLR, suggesting that stereo sounds are useful training signals. Augmentations are important for all the models. We also note that the data distribution between training and test cases is quite different (i.e., training with music signals and testing on the human speech), suggesting that our approach is capable of generalization.

Next, we ask how well our time delay estimation methods can localize sound directions in challenging real-world scenes, using audio collected from the internet.

We collected 30 internet binaural videos, extracted 1K samples from them, and use human judgments to label sound directions. These videos contain a variety of sounds, including engine noise and human speech, which are often far from the viewer. Many also contain multiple sound sources and background noise. We provide examples of these videos on our project webpage.

Since it is difficult for humans to describe sound directions in terms of time delay, we asked listeners to annotate the direction of the loudest sound. The annotator (one of the authors) listened to the audio with headphones and labeled 5 directions: left/right, center left/right, and center. From these, we created binary left/right labels, which can be objectively evaluated by thresholding the delay: we discard the center label and merge the remaining directional labels (resulting in 885 examples). We measure the accuracy of the thresholded time delay using these labels. We balance the dataset by swapping stereo channels, such that chance is 50%.In the original version of the paper, we evaluated on the raw (unbalanced) data. Since arXiv v3, we balanced the dataset by swapping stereo channels (Tab. 4.2 and 9). As in the mixture experiments, we provide models with 0.5s audio to ensure that they have sufficient context. We also compared with a method that uses interaural intensity difference (IID) cues, by comparing the root mean square (RMS) of each audio channel to determine which channel is louder to predict its left/right direction (equivalent to thresholding based on ∣∣x∣∣||{\mathbf{x}}||).

To help understand how our predicted time delays vary with motion and change over time, we correlated the visual motions in our dataset with the predicted time delays. We tracked race cars in the subset of our in-the-wild dataset that contains them using CenterTrack , and manually removed erroneous tracks (obtaining 49 trajectories). We applied our StereoCRW model (with 1024 samples and 128 votes). In qualitative examples (Fig. 4), we see that the car’s on-screen position is closely correlated with the time delay. Interestingly, our model continues to convey the car’s position when it moves off-screen. We computed the Spearman rank correlation coefficient between the xx position of each track (using the center of the bounding box) and the time delay predictions (averaged over all video clips), yielding ρ=0.57\rho=0.57 for our model and ρ=0.48\rho=0.48 for GCC-PHAT. More video results can be found on our project webpage.

We also found that our model worked successfully on sounds from ordinary video recordings from recent iPhones. For qualitative results, please see Sec. A.1.

Instead of attending to a loud, dominant sound in a mixture (Sec. 4.1), we use vision to specify which, of several, sound sources to localize. The model solves this task by first learning to associate a voice with the given visual attribute of the speakers. Unlike audio spatialization and active speaker detection tasks, which provide temporal and spatial cues from the video, the audio-visual streams in our case are not necessarily temporally aligned (e.g., modeling the challenges of tracking a speaker when they move out of sight).

We evaluate the audio-visual model in a simulated environment. This allows us to control the positions of the speakers, and to remove other localization cues from the images and audio. We use audio clips from VoxCeleb2 with the simulation parameters from Sec. 4.1. We select 500 speakers from the database (same as training) and pair them with their corresponding face images. Following the work in source separation, we remove loudness cues by normalizing the volume of sound sources, and placing two speakers in the simulator at the same distance (but at different angles). Note that, since the position of the speakers is randomized, the visual signal does not provide localization cues (e.g., via perspective). We also convert delay predictions to direction-of-arrival angles, using the known radius and microphone distance.

Comparisons.

To provide points of comparison for this novel task, we compare our audio-visual approach with audio-only methods including GCC-PHAT. We also provide a (oracle) baseline which selects one of the two speakers’ ground-truth time delay at random, thus simulating a method that perfectly solves the localization task but which is unable to match faces to voices. We also consider a two-stage method that first separates the speaker’s voice for each channel using VisualVoice , a state-of-the-art audio-visual separation model, then applies audio-based time delay estimation methods to the separated sound. To ensure a fair comparison, we retrain the static-image based separation model with input audio duration of 1.27s and 2.55s.

We perform visually-guided time delay estimation on a self-recorded video (Fig. 7). Two speakers talk concurrently while moving off-screen. Our model localizes each speaker in the mixture with a cropped image of their face. We show the mean and standard deviation of delay predictions in 2.0s windows.

We have proposed to use self-supervised time delay estimation to localize sounds by learning interaural correspondence. We also introduced a novel visually-guided localization task. Our audio models obtain performance on par with supervised methods on real-world sound, while our audio-visual model successfully localizes speakers in mixtures. We see our work opening two directions: first, integrating more visual information for multisensory localization and, second, finding finer-grained delays using recent methods from optical flow .

Our audio-visual model associates the appearance of speakers with the sound of their voice. We tested on speakers that our model has been trained on, avoiding the need to generalize based solely on a person’s appearance . However, there is still a potential for the model to exhibit bias. Our sound-based models are trained on music, which may not be representative of all downstream tasks. We released code, data, and models on our project site.

Acknowledgments.

We would like to thank Xixi Hu for her valuable suggestions about the augmentation idea on random walk graphs. We thank Justin Salamon for helpful discussions and Daniele Salvati for the help on the simulator setup. We thank Zhaoying Pan and Matthew Sticha for the help recording real-world examples. We also thank Daniel Geng for his comments and feedback on the paper. This work was funded in part by DARPA Semafor and Cisco Systems. The views, opinions and/or findings expressed are those of the authors and should not be interpreted as representing the official views or policies of the Department of Defense or the U.S. Government.

Appendix A.1 Qualitative Results on Phone Recordings

We also ran our model on ordinary iPhone-recorded videos, exploiting the fact that portrait-mode recordings have a sufficiently large baseline for estimating time delays. We use a combination of self-collected videos and internet videos (we used an iPhone 12 for self-recorded videos and searched Flickr for videos with tags indicating that they were recorded with an iPhone 13). We provide qualitative results in Fig. 6. Please see our webpage for more video results.

Appendix A.2 Ablation Study

We study the effect of the number of votes, mm, used during post-processing for both StereoCRW and GCC-PHAT. We evaluate 1024-sample audio using m∈{1,32,128,256,512}m\in\{1,32,128,256,512\} for both mean and mode. In the special case m=1m=1, the result is not affected by post-processing, and purely measures the quality of the representation. As shown in Fig. 7(a), both methods improve with the number of votes. Our method benefits from mode post-processing, while mean works better for GCC-PHAT. In particular, we significantly outperform GCC-PHAT with m=1m=1 vote, emphasizing the quality of our representation.

We also design experiments to study the performance gap between using probabilistic post-processing and nearest neighbor (argmax) for our method. For consistency, we evaluated our StereoCRW model with both types of post-processing on two simulated environments, with different vote numbers. The results are shown in Fig. 8. Probabilistic post-processing outperforms the nearest neighbor post-processing in the complex simulated environment. We use probabilistic post-processing in our other experiments.

Duration.

We ask how our model performs when given longer audio, exploiting the fact that our embeddings use fully convolutional networks and thus can be tested on arbitrary-sized inputs (Fig. 7(b)). We provided our method with various input sizes (up to 8×8\times the training duration). For very long audio (4×4\times training duration), we found that our model’s performance starts saturating, and that GCC-PHAT overtakes it. This may be due to the fact that the model has a fixed-size (d=128d=128) representation, while GCC-PHAT grows its representation—the waveform itself—with the input size and eventually converges to the correct solution (a maximum likelihood estimate in many situations).

Simulated vs. real data.

We also study the data distribution gap between simulation and real-world data. We trained our best self-supervised model (StereoCRW) and Salvati et al. on the simulated data with Free-Music-Archive clips. We evaluated on both simulated and in-the-wild recordings (Tab. 9). As expected, our model trained on FMA-Sim obtains competitive performance, but overall does not perform as well as a model trained on real data. The supervised model improves on the in-the-wild evaluation while the performance drops on the simulated evaluation cases when training on FMA-Sim. We also include the comparison between mode and mean post-process for Salvati et al. in the Tab. 9.

Appendix A.3 Training with Youtube-ASMR

We use a ResNet with 9 layers as the backbone for the audio encoder. We modify the input channel number of the first convolution layer to be 2, and the output of the last fully-connected layer as 128. For a raw waveform of length LL, we use an STFT with a window length of 256 and hop length of ⌊L128⌋\lfloor\frac{L}{128}\rfloor to create an input spectrogram. For the audio-visual task, we use a hop length of 160 to create an input spectrogram.

Augmentations.

During the training, we apply the following augmentations to audio where the first three are regular augmentations applied to all the models and the last two are applied to augmented models only:

Random channel swapping: we randomly swap the left and right audio channels with a probability of 0.5.

Random channel-wise scaling: we randomly re-scale each audio channel by the factor in the range of [0.5,1.5][0.5,1.5].

Random shifting: for the instance discrimination model with mono audio, we randomly shift the audio for −16-16 to 1616 samples. For the audio-visual model, we apply different random shifts for each mono audio with −24-24 to 2424 samples.

Random noise: we add random Gaussian noise to audio with SNR=$$.

Random reverberation: we add random reverberation to audio with RT60=[0,0.9][0,0.9].

Mixture augmentation: we randomly add another sound to the original audio. We normalize the second sound to be 10% – 100% loudness of the original audio before mixing. For the audio-visual model, the second sound is normalized to be 50%-150% intensity level of the original one.

When computing affinity matrix A21A_{21} for the contrastive random walk model, we do not augment x1{\mathbf{x}}_{1} with noise or mixture augmentation, so as to avoid learning unexpected matching. Similarly, for the instance discrimination model, we do not apply noise, reverberate or mixture augmentations to one of the two channels.

Training details.

To accelerate the training process, we first train each model with 0.48s audio (7680 samples) and then finetune on the input audio of 0.064s (1024 samples) using a correspondingly finer hop length.