BatVision: Learning to See 3D Spatial Layout with Two Ears

Jesper Haahr Christensen, Sascha Hornauer, Stella Yu

I Introduction

Our task is to train a machine learning system that can turn binaural sound signals to visual scenes. Solving this challenge would benefit robot navigation and machine vision, especially in low-light or no-light conditions.

While many animals sense the spatial layout of the world through vision, some species such as bats, dolphins, and whales rely heavily on acoustic information. For example, bats have advanced ears that give them a form of vision in the dark known as echolocation: They sense the world by continuously emitting ultrasonic pulses and processing echos returned from the environment.

It is indeed possible to locate highly reflecting ultrasonic targets in the 3D space by using an artificial pinnae pair of bats, which acts as complex direction dependent spectral filters and using head related transfer functions .

Likewise, humans who suffer from vision loss have shown to develop capabilities of echolocation using palatal clicks similar to dolphins, learning to sense obstacles in the 3D space by listening to the returning echoes .

Inspired by bats’ echolocation, we design BatVision that can form a visual image of the 3D world by just listening to the environmental echo sound with two ears (Fig. 1).

Contrary to existing works , our system uses only two simple low-cost consumer-grade microphones to keep it small, mobile, and easily reproducible. Our microphones are embedded into a human pinnae pair to utilize the spectral filters of an emulated human auditory system, which has an additional benefit of easy debugging by human engineers.

Mounted on a model car, our BatVision also has a speaker and a camera which is only used during training for providing visual image ground-truth. Like bats, our speaker emits frequency modulated chirps in the audible spectrum, and our microphones receive echos returned from the environment. Our camera captures stereo image pairs of the scene ahead, from which depth disparity maps can be calculated.

During training, we first collect a dataset of time-synchronized binaural audio signals and stereo image pairs in an indoor office environment, and then train a neural network model to predict images such as depth maps and grayscale images from audio data alone.

During testing, we just need the sound signals to reconstruct depth maps or grayscale images. By just listening with two ears, which receive sound echos at only two points in the 3D space, our BatVision is able to generate a depth map of the 3D space ahead that resolves features such as walls, hallways, door openings, and roughly outlined furniture correctly in azimuth, elevation, and distance, whereas our reconstructed grayscale images show surprisingly plausible floor layouts even though obstacles lack finer details.

For a navigation system, such an intelligent sound system could provide information complementary to vision sensors, independent of light and at very low additional costs. Our approach is conceptually simple, practically easy to implement, and readily deployable on embedded mobile platforms.

To the best of our knowledge, our BatVision is the first work that generates scene depth maps from binaural sound only. Our code, model, and data are available at https://github.com/SaschaHornauer/Batvision.

II Related Works

In , microphones are placed in an artificial bat pinnae to receive the sound signal. The natural form of the bat pinnae acts as a frequency filter, useful for separating spatial information in both azimuth and elevation . These works motivate our use of short FM chirps and artificial human pinnaes with integrated microphones.

In , the task is to recognize scenes from echo-cochleogram fingerprints and to create a topological map of the surrounding. In , the goal is to autonomously drive a mobile robot while mapping and avoiding obstacles using azimuth and range information from ultrasonic sensors. They classify echo spectrograms into obstacles or not, biological objects or not, along a single scan line and without visual reconstruction of the scene.

In , ultrasonic echoes are recorded, dilated, and played back to a human subject in the audible spectrum. After initial training, human subjects were able to pick up echolocation abilities to estimate azimuth, distance, and to some extent, elevation of targets. In , 3D targets are localized based on an array of microphones instead of binaural microphones.

Sound Source Localization. In , deep neural network models are trained to localize the source of the sound (e.g. a piano) in images or videos. Remarkable results are obtained in a self-supervised learning framework, demonstrating the potential of learning associations between paired audio-visual data.

In , sound is localized using an acoustic camera , a hybrid audio-visual sensor that provides RGB video overlaid with acoustic sound, aligned in time and space. All the works on sound localization receive sound signals passively.

In , an audio monologue of a speaker is turned into visual gestures of the speaker’s arms and hands, by translating audio clips into 2D2D trajectories of joint coordinates. Our sound to vision decoder model is inspired by their cross-domain translation success.

III Collection of Our Audio-Visual Dataset

We collect a new dataset of time-synchronized binaural audio, RGB images, and depth maps, which can be used for learning the associations between sound and vision.

We traverse an office building in the hallways, open areas, conference rooms, and office spaces. We fix our BatVision on a trolley and slowly push it around, so that there is no active motor noise corrupting our sound.

We collect data at various spatial locations to minimize correlation and maximize scene diversity (Fig 2). A total of 39,500 and 7,500 instances at two different parts of the same floor are collected for training and validation respectively, with additional 5,040 instances on a different floor for testing. While hallways appear similar, their spatial layout, furniture, occupancy, and decorations are different.

III-B Our Hardware: Speaker, Ears, and Camera

We use a ZED camera to capture stereo image pairs and extract depth maps from them. Our camera, speaker, and artificial ears are mounted on a small model car (Fig. 1).

III-C Our Audio clips and Visual Images

We synchronize all the audio instances by the time of the recorded chirp. However, during training, we augment the audio data by jittering the position of the window by 30%.

IV Our Sound to Vision Prediction Models

We use an encoder-decoder network architecture to turn the audio clip into the visual image, and further improve the quality of generated images using an adversarial discriminator to contrast them against the ground-truth (Fig. 4).

We train our model with two possible audio representations. Our experiments indicate that spectrograms yield slightly better sound-to-vision predictions over raw waveforms. However, as we aim for a real-time BatVision system on embedded platforms, we focus on raw waveforms which are more computationally efficient.

Encoder for Waveforms. Following SoundNet , we represent the binaural input as two channels of 1D1D signals and transform it into a 1024-dimensional feature vector with 8 temporal convolutions. See Fig. 5 and Table I for details.

Encoder for Spectrograms. Likewise, with successive temporal convolutions and downsampling, we gradually reduce the time-dimension of the spectrograms down to 1, producing a 1×f×10241\times f\times 1024 feature vector, where ff is the number of final frequencies. ff depends on the downsampling factors along the y-axis of the spectrogram.

IV-B Our Visual Image Generator G𝐺G

The generator decodes the latent audio feature vector and expands it into visual scene image. For raw waveforms, successive deconvolutions yield the best results, whereas for spectrograms, a UNet-type encoder-decoder network yields best results. We investigate several resolutions for reconstructed images, from 16×1616\times 16 to 128×128128\times 128.

Decode by A UNet. To transform the output of our audio encoder to a 2D2D image representation suitable for a UNet, we reshape the 1024-dimensional feature vector into a 32×32×132\times 32\times 1 tensor. For spectrograms, where the audio encoder outputs a 1×f×10241\times f\times 1024 vector and f≠1f\neq 1, we first apply two fully connected linear layers before reshaping it into a 32×32×132\times 32\times 1 tensor. The output of this generator depends on the target resolution, e.g. 128×128×1128\times 128\times 1.

The encoder of the UNet downsamples the 32×32×132\times 32\times 1 input through several layers of double convolutions followed by batch normalization and ReLU, whereas the decoder of the UNet upsamples the input through double de-convolutions followed by batch normalization and ReLU. Skip connections are utilized wherever possible.

Decode from Direct Upsampling. Given the 1×1×10241\times 1\times 1024 latent audio vector, we apply a series of upsampling layers (as in the UNet decoder) to reach the target resolution. See the layer configuration for the 128×128×1128\times 128\times 1 output in Table II.

IV-C Our Adversarial Discriminator

We add an adversarial discriminator DD for generating more detailed and realistic predictions. We implement the discriminator as a PatchGAN to ensure that the predicted visual image has similar looking patches as the set of ground-truth images; DD tries to classify whether each N×NN\times N patch looks real or fake as a ground truth sample, where NN is roughly 1/31/3 of the image size. DD consists of a few convolutional layers with depth, kernel size and stride parameters dependent on the final output image size. See the layer configuration in Table III for size 128×128128\times 128.

V Experimental Results

In a preliminary study that compares input modes and fusion design choices, we predict small images at size 16×1616\times 16. We have the following observations.

For raw waveforms, early fusion (left-right-channel concatenation of the input audio) outperforms late fusion (concatenation at Conv8, see Fig. 5).

Spectrograms yield slightly better results than raw waveforms.

However, as we aim for real-time performance on embedded platforms, we focus on the least computationally expensive method using waveforms.

We compute the prediction error via an L1L_{1} regression loss:

where xx is the audio waveforms or spectrograms, yy is the ground truth visual image (depth map or grayscale scene image), AA is the audio encoder, and GG is the generator.

We use leaky ReLU with slope 0.20.2, batch size 1616, and Adam solver with an initial learning rate of 1×10−41\times 10^{-4} with parameters β1\beta_{1} and β2\beta_{2} set to 0.90.9 and 0.9990.999 respectively.

Table IV compares various model choices along with two trivial reconstruction baselines which do not learn any sound and vision associations at all:

Random uniform noise in the [0,1)\left[0,1\right) range.

For raw waveforms, direct upsampling and early fusion perform the best. For spectrograms, early fusion, downsampling to 1×10×10241\times 10\times 1024 and the UNet generator perform best. These two best configurations are retrained for output dimensions of 32×3232\times 32, 64×6464\times 64 and 128×128128\times 128, and the loss is higher for a larger depth map (Table V).

Fig. 6 compares reconstructions at different resolutions. Fig. 7 shows more samples of diverse scenes at reconstruction size 128×128128\times 128. The sound-to-vision predictions provide a rough outline of the spatial layout of the 3D scene.

V-B Generator with Adversarial Discriminator

We use an Generative Adversarial network (GAN) model at the patch level to improve the visual reconstruction quality. We use the following least-squares loss instead of a sigmoid cross-entropy loss in order to avoid vanishing gradients :

where λ\lambda is a weight factor. We use leaky ReLU with slope 0.2, λ=100\lambda=100, batch size 1616, and Adam solver with learning rate set to 2×10−42\times 10^{-4} with parameters β1\beta_{1} and β2\beta_{2} set to 0.50.5 and 0.9990.999 respectively.

Table V compares the test set loss over a few design choices. As in the ”Generator Only” case, the loss is moderately higher for a larger depth map. However, Fig. 7 shows our sample reconstructions by GAN have much finer details and clearer borders, and our grayscale reconstructions in the rightmost column have well placed floors even though objects are roughly outlined and abstracted.

V-C Limitations of Our Approach

How sound resonates, propagates and reflects in a room has a huge impact on sound-to-vision predictions.

Some materials have dampening properties, leading to faint or absorbed echos.

Facing corners, where hallways fork in different directions, poses a big challenge, because sound waves scatter off in different directions.

In areas with dense obstacles such as conference rooms with many office chairs, our sound-to-vision model often fails to predict any meaningful content (Fig. 8).

VI Conclusions

Our BatVision system with a trained sound-to-vision model can reconstruct depth maps from binaural sound recorded by only two microphones to a remarkable accuracy. It can predict detailed indoor scene depth and obstacles such as walls and furniture. Sometimes, it even outperforms our ground-truth depth map obtained from a stereo vision algorithm which struggles to estimate disparity reliably.

Generating the grayscale scene image is more difficult; the amount of detail and information required is not expected to be present in sound echos. However, our trained model is able to generate plausible wall placements and free floor areas. When objects are not recognizable from the sound, the network fills in with an approximation of obstacles.

Such seemingly incredible sound-to-vision results reflect natural statistical correlations between the sound and the image of indoor scenes, captured by our model trained on diverse scenes and likely utilized in a similar fashion by humans and animals.

References