BatVision: Learning to See 3D Spatial Layout with Two Ears
Jesper Haahr Christensen, Sascha Hornauer, Stella Yu
I Introduction
Our task is to train a machine learning system that can turn binaural sound signals to visual scenes. Solving this challenge would benefit robot navigation and machine vision, especially in low-light or no-light conditions.
While many animals sense the spatial layout of the world through vision, some species such as bats, dolphins, and whales rely heavily on acoustic information. For example, bats have advanced ears that give them a form of vision in the dark known as echolocation: They sense the world by continuously emitting ultrasonic pulses and processing echos returned from the environment.
It is indeed possible to locate highly reflecting ultrasonic targets in the 3D space by using an artificial pinnae pair of bats, which acts as complex direction dependent spectral filters and using head related transfer functions .
Likewise, humans who suffer from vision loss have shown to develop capabilities of echolocation using palatal clicks similar to dolphins, learning to sense obstacles in the 3D space by listening to the returning echoes .
Inspired by bats’ echolocation, we design BatVision that can form a visual image of the 3D world by just listening to the environmental echo sound with two ears (Fig. 1).
Contrary to existing works , our system uses only two simple low-cost consumer-grade microphones to keep it small, mobile, and easily reproducible. Our microphones are embedded into a human pinnae pair to utilize the spectral filters of an emulated human auditory system, which has an additional benefit of easy debugging by human engineers.
Mounted on a model car, our BatVision also has a speaker and a camera which is only used during training for providing visual image ground-truth. Like bats, our speaker emits frequency modulated chirps in the audible spectrum, and our microphones receive echos returned from the environment. Our camera captures stereo image pairs of the scene ahead, from which depth disparity maps can be calculated.
During training, we first collect a dataset of time-synchronized binaural audio signals and stereo image pairs in an indoor office environment, and then train a neural network model to predict images such as depth maps and grayscale images from audio data alone.
During testing, we just need the sound signals to reconstruct depth maps or grayscale images. By just listening with two ears, which receive sound echos at only two points in the 3D space, our BatVision is able to generate a depth map of the 3D space ahead that resolves features such as walls, hallways, door openings, and roughly outlined furniture correctly in azimuth, elevation, and distance, whereas our reconstructed grayscale images show surprisingly plausible floor layouts even though obstacles lack finer details.
For a navigation system, such an intelligent sound system could provide information complementary to vision sensors, independent of light and at very low additional costs. Our approach is conceptually simple, practically easy to implement, and readily deployable on embedded mobile platforms.
To the best of our knowledge, our BatVision is the first work that generates scene depth maps from binaural sound only. Our code, model, and data are available at https://github.com/SaschaHornauer/Batvision.
II Related Works
In , microphones are placed in an artificial bat pinnae to receive the sound signal. The natural form of the bat pinnae acts as a frequency filter, useful for separating spatial information in both azimuth and elevation . These works motivate our use of short FM chirps and artificial human pinnaes with integrated microphones.
In , the task is to recognize scenes from echo-cochleogram fingerprints and to create a topological map of the surrounding. In , the goal is to autonomously drive a mobile robot while mapping and avoiding obstacles using azimuth and range information from ultrasonic sensors. They classify echo spectrograms into obstacles or not, biological objects or not, along a single scan line and without visual reconstruction of the scene.
In , ultrasonic echoes are recorded, dilated, and played back to a human subject in the audible spectrum. After initial training, human subjects were able to pick up echolocation abilities to estimate azimuth, distance, and to some extent, elevation of targets. In , 3D targets are localized based on an array of microphones instead of binaural microphones.
Sound Source Localization. In , deep neural network models are trained to localize the source of the sound (e.g. a piano) in images or videos. Remarkable results are obtained in a self-supervised learning framework, demonstrating the potential of learning associations between paired audio-visual data.
In , sound is localized using an acoustic camera , a hybrid audio-visual sensor that provides RGB video overlaid with acoustic sound, aligned in time and space. All the works on sound localization receive sound signals passively.
In , an audio monologue of a speaker is turned into visual gestures of the speaker’s arms and hands, by translating audio clips into trajectories of joint coordinates. Our sound to vision decoder model is inspired by their cross-domain translation success.
III Collection of Our Audio-Visual Dataset
We collect a new dataset of time-synchronized binaural audio, RGB images, and depth maps, which can be used for learning the associations between sound and vision.
We traverse an office building in the hallways, open areas, conference rooms, and office spaces. We fix our BatVision on a trolley and slowly push it around, so that there is no active motor noise corrupting our sound.
We collect data at various spatial locations to minimize correlation and maximize scene diversity (Fig 2). A total of 39,500 and 7,500 instances at two different parts of the same floor are collected for training and validation respectively, with additional 5,040 instances on a different floor for testing. While hallways appear similar, their spatial layout, furniture, occupancy, and decorations are different.
III-B Our Hardware: Speaker, Ears, and Camera
We use a ZED camera to capture stereo image pairs and extract depth maps from them. Our camera, speaker, and artificial ears are mounted on a small model car (Fig. 1).
III-C Our Audio clips and Visual Images
We synchronize all the audio instances by the time of the recorded chirp. However, during training, we augment the audio data by jittering the position of the window by 30%.
IV Our Sound to Vision Prediction Models
We use an encoder-decoder network architecture to turn the audio clip into the visual image, and further improve the quality of generated images using an adversarial discriminator to contrast them against the ground-truth (Fig. 4).
We train our model with two possible audio representations. Our experiments indicate that spectrograms yield slightly better sound-to-vision predictions over raw waveforms. However, as we aim for a real-time BatVision system on embedded platforms, we focus on raw waveforms which are more computationally efficient.
Encoder for Waveforms. Following SoundNet , we represent the binaural input as two channels of signals and transform it into a 1024-dimensional feature vector with 8 temporal convolutions. See Fig. 5 and Table I for details.
Encoder for Spectrograms. Likewise, with successive temporal convolutions and downsampling, we gradually reduce the time-dimension of the spectrograms down to 1, producing a feature vector, where is the number of final frequencies. depends on the downsampling factors along the y-axis of the spectrogram.
IV-B Our Visual Image Generator G𝐺G
The generator decodes the latent audio feature vector and expands it into visual scene image. For raw waveforms, successive deconvolutions yield the best results, whereas for spectrograms, a UNet-type encoder-decoder network yields best results. We investigate several resolutions for reconstructed images, from to .
Decode by A UNet. To transform the output of our audio encoder to a image representation suitable for a UNet, we reshape the 1024-dimensional feature vector into a tensor. For spectrograms, where the audio encoder outputs a vector and , we first apply two fully connected linear layers before reshaping it into a tensor. The output of this generator depends on the target resolution, e.g. .
The encoder of the UNet downsamples the input through several layers of double convolutions followed by batch normalization and ReLU, whereas the decoder of the UNet upsamples the input through double de-convolutions followed by batch normalization and ReLU. Skip connections are utilized wherever possible.
Decode from Direct Upsampling. Given the latent audio vector, we apply a series of upsampling layers (as in the UNet decoder) to reach the target resolution. See the layer configuration for the output in Table II.
IV-C Our Adversarial Discriminator
We add an adversarial discriminator for generating more detailed and realistic predictions. We implement the discriminator as a PatchGAN to ensure that the predicted visual image has similar looking patches as the set of ground-truth images; tries to classify whether each patch looks real or fake as a ground truth sample, where is roughly of the image size. consists of a few convolutional layers with depth, kernel size and stride parameters dependent on the final output image size. See the layer configuration in Table III for size .
V Experimental Results
In a preliminary study that compares input modes and fusion design choices, we predict small images at size . We have the following observations.
For raw waveforms, early fusion (left-right-channel concatenation of the input audio) outperforms late fusion (concatenation at Conv8, see Fig. 5).
Spectrograms yield slightly better results than raw waveforms.
However, as we aim for real-time performance on embedded platforms, we focus on the least computationally expensive method using waveforms.
We compute the prediction error via an regression loss:
where is the audio waveforms or spectrograms, is the ground truth visual image (depth map or grayscale scene image), is the audio encoder, and is the generator.
We use leaky ReLU with slope , batch size , and Adam solver with an initial learning rate of with parameters and set to and respectively.
Table IV compares various model choices along with two trivial reconstruction baselines which do not learn any sound and vision associations at all:
Random uniform noise in the range.
For raw waveforms, direct upsampling and early fusion perform the best. For spectrograms, early fusion, downsampling to and the UNet generator perform best. These two best configurations are retrained for output dimensions of , and , and the loss is higher for a larger depth map (Table V).
Fig. 6 compares reconstructions at different resolutions. Fig. 7 shows more samples of diverse scenes at reconstruction size . The sound-to-vision predictions provide a rough outline of the spatial layout of the 3D scene.
V-B Generator with Adversarial Discriminator
We use an Generative Adversarial network (GAN) model at the patch level to improve the visual reconstruction quality. We use the following least-squares loss instead of a sigmoid cross-entropy loss in order to avoid vanishing gradients :
where is a weight factor. We use leaky ReLU with slope 0.2, , batch size , and Adam solver with learning rate set to with parameters and set to and respectively.
Table V compares the test set loss over a few design choices. As in the ”Generator Only” case, the loss is moderately higher for a larger depth map. However, Fig. 7 shows our sample reconstructions by GAN have much finer details and clearer borders, and our grayscale reconstructions in the rightmost column have well placed floors even though objects are roughly outlined and abstracted.
V-C Limitations of Our Approach
How sound resonates, propagates and reflects in a room has a huge impact on sound-to-vision predictions.
Some materials have dampening properties, leading to faint or absorbed echos.
Facing corners, where hallways fork in different directions, poses a big challenge, because sound waves scatter off in different directions.
In areas with dense obstacles such as conference rooms with many office chairs, our sound-to-vision model often fails to predict any meaningful content (Fig. 8).
VI Conclusions
Our BatVision system with a trained sound-to-vision model can reconstruct depth maps from binaural sound recorded by only two microphones to a remarkable accuracy. It can predict detailed indoor scene depth and obstacles such as walls and furniture. Sometimes, it even outperforms our ground-truth depth map obtained from a stereo vision algorithm which struggles to estimate disparity reliably.
Generating the grayscale scene image is more difficult; the amount of detail and information required is not expected to be present in sound echos. However, our trained model is able to generate plausible wall placements and free floor areas. When objects are not recognizable from the sound, the network fills in with an approximation of obstacles.
Such seemingly incredible sound-to-vision results reflect natural statistical correlations between the sound and the image of indoor scenes, captured by our model trained on diverse scenes and likely utilized in a similar fashion by humans and animals.