AutoAD: Movie Description in Context

Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, Andrew Zisserman

Introduction

That of all the arts, the most important for us is the cinema. Vladimir Lenin

One of the long-term aims of computer vision is to understand long-form feature films. There has been steady progress towards this aim with the identification of characters by their face and voice , the recognition of their actions and inter-actions , of their relationships , and 3D pose . However, this is still a long way away from story understanding. Movie Audio Description (AD), the narration describing visual elements in movies, provides a means to evaluate current movie understanding capabilities. AD was developed to aid visually impaired audiences, and is typically generated by experienced annotators. The amount of AD on the internet is growing due to more societal support for visually impaired communities and its inclusion is becoming an emerging legal requirement.

AD differs from image or video captioning in several significant respects , bringing its own challenges. First, AD provides dense descriptions of important visual elements over time. Second, AD is always provided on a separate soundtrack to the original audio track and is highly complementary to it. It is complementary in two ways: it does not need to provide descriptions of events that can be understood from the soundtrack alone (such as dialogue and ambient sounds), and it is constrained in time to intervals that do not overlap with the dialogue. Third, unlike dense video captioning, AD aims at storytelling; therefore, it typically includes factors like a character’s name, emotion, and action descriptions.

In this work, our objective is automatic AD generation – a model that takes continuous movie frames as input and outputs AD in text form. Specifically, we generate text given a temporal interval of an AD, and evaluate its quality by comparing with the ground-truth AD. This is a relatively unexplored task in the vision community with previous work targeting ActivityNet videos , a very different domain to long-term feature films with storylines, and the LSMDC challenge , where the descriptions and character names are treated separately.

As usual, one of the challenges holding back progress is the lack of suitable training data. Paired image-text or video-text data that is available at scale, such as alt-text or stock footage with captions , does not generalize well to the movie domain . However, collecting high-quality data for movie understanding is also difficult. Researchers have tried to hire human annotators to describe video clips but this does not scale well. Movie scripts, books and plots have also been used as learning signals but they do not ground on vision closely and are limited in number.

In this paper we address the AD and training data challenges by – Spoiler Alert – developing a model that uses temporal context together with a visually conditioned generative language model, while providing new and cleaner sources of training data. To achieve this, we leverage the strength of large-scale language models (LLMs), like GPT , and vision-language models, like CLIP , and integrate them into a video captioning pipeline that can be effectively trained with AD data.

Our contributions are the following: (i) inspired by ClipCap we propose a model that is effectively able to leverage both temporal context (from previously generated AD) and dialogue context (in particular the names of characters) to improve AD generation. This is done by bridging foundation models with lightweight adapters to integrate both types of context; (ii) we address the lack of large-scale training data for AD by pretraining components of our model on partially missing data which are typically available in large quantities e.g. text-only AD without movie frames, or visual captioning datasets without multiple sentences as context; (iii) we propose an automatic pipeline for collecting AD narrations at scale using speaker-based separation; and finally (iv) we show promising results on automatic AD, as seen from both qualitative and quantitative evaluations, and also achieve impressive zero-shot results on the LSMDC multi-description benchmark comparable to the finetuned state-of-the-art.

Related Works

Image Captioning. Image captioning is a long-standing problem in computer vision . Early pioneering works learn to associate images and words within a limited vocabulary and a set of images . Large-scale image captioning datasets have been collected by scraping images from the internet and their corresponding alt-texts with quality filters as a post-processing . In doing so, strong joint image-text representations can be learned , and image captioning from raw pixels, with impressive results . Recent work learns a bridge between strong joint image-text representations (CLIP) and the natural language representation (GPT-2) for image captioning, obtaining promising results that generalise well across domains. In this work, we extend this approach to perform automatic AD from videos.

Video Captioning. Video captioning presents additional challenges due to the lack of quality large-scale video-text data and increased complexity from the temporal axis. Early video caption datasets adopt manual annotations, a far from scalable collection method. ASR (automated speech recognition) from YouTube instructional videos is collected at scale for video-language datasets , but contains high levels of noise due to the weak correspondence between the narration and visual content. VideoCC transfers captions from images to videos, but this method is still limited by the existing seed image captioning dataset used. Earlier video captioning models lack generalisation capabilities due to limited training data . Some recent methods train on ASR from the HowTo100M dataset, while others expand image-text representations to multiple frames.

A task more related to AD is that of dense video captioning , which involves producing a number of captions and their corresponding grounded timestamps in the video. To enrich inter-task interactions, recent works for this task jointly train both a captioning and localization module. Our task differs in that the captions are: made with the intent to aid storytelling; specific to the movie domain; and complementary to the audio track.

Visual Storytelling. Most similar in vein to the AD task is visual storytelling , in which the goal is to generate coherent sentences for a sequence of video clips or images. LSMDC proposes the multi-description task of generating captions for a set of clips from a movie, with character names anonymized. In contrast, movie AD takes as input a continuous long video and describes the visual happenings complementary to the story, characters, dialogue and audio. Most similar to our model is TPAM which prompts a frozen GPT-2 with local visual features. Ours differs in that: (i) it is not restricted to local visual context but rather global by recurrently conditioning on previous outputs; and (ii) we additionally pretrain GPT on in-domain text-only AD data.

Movie Understanding. Previous works investigate storyline understanding by aligning movies to additional data sources such as plots , books , scripts , and YouTube summaries . However, these sources are limited in number and often do not closely relate to the visual elements in the frame. Using existing movie AD as the data source for videos is an emergent direction for movie understanding. LSMDC , M-VAD dataset and MPII-MD , gather AD and scripts from movies to provide captions for short video clips, several seconds in duration. QuerYD provides high-quality textual descriptions for longer videos by scraping AD from YouDescribe , an online community of AD contributors. Recently, the MAD dataset collects movie AD at scale to provide dense textual annotations for movies with a focus on visual grounding task.

Prompt Tuning and Adapters. Originally for language modelling, prompt tuning is a lightweight approach to adapt a pretrained model to perform a downstream task. Early works learn prompt vectors that are shared within the targeted dataset and task. A similar line of works to ours is visual-conditioned prompt tuning, in which the prompt vectors are conditioned on the visual inputs. Visual-conditioned prompts are used for adapting pretrained image-language models , and for few-shot learning . Training lightweight feature adapters between pretrained vision and text encoders is another approach to adapt pretrained models . The adapter layers can also be inserted into the pretrained language model in an interleaved way . Our work adopts prompt tuning in order to condition a language generation model on visual information (frames), and textual context (subtitles and previous AD).

Method

Given a long-form movie V\mathcal{V} segmented into multiple short clips {x1,x2,...,xT}\{\mathbf{x}_{1},\mathbf{x}_{2},...,\mathbf{x}_{T}\}, our goal is to generate the audio description (AD) in text form for every movie clip. Note that each movie clip is cut from the raw movie based on the timestamp [tstart,tend][t_{\text{start}},t_{\text{end}}] given by the AD annotation. Specifically, for the ii-th movie clip consisting of multiple frames xi={I1,I2,...,IN}\mathbf{x}_{i}=\{\mathcal{I}_{1},\mathcal{I}_{2},...,\mathcal{I}_{N}\}, we aim to produce text Ti\mathcal{T}_{i} that describes the visual elements in such a way that helps the visually impaired follow the storyline. To this purpose, an ideal AD generation system must be able to exploit the full contextual information leading up to the ii-th movie clip. One method for this, which we adopt, is to use previous AD Tt<i\mathcal{T}_{t<i} and subtitles St≤i\mathcal{S}_{t\leq i} to generate the text Ti\mathcal{T}_{i}. In the following sections, we first give an overview of our visual captioning pipeline with prompt tuning (Sec. 3.1), followed by our contextual components (Sec. 3.2), and finally the pretraining methods with partial data (Sec. 3.3).

In order to describe our method, we first present the typical pipeline for an image captioning model, and then detail how we extend this to ingest multiple frames and additional text context. Given an image-caption pair {Ii,Ci}\{\mathcal{I}_{i},\mathcal{C}_{i}\}, where the caption consists of a sequence of language tokens Ci={c1,c2,...,ck}\mathcal{C}_{i}=\{c_{1},c_{2},...,c_{k}\}, the standard objective of an image captioning model is to generate text tokens C^i\hat{\mathcal{C}}_{i} that are close to the target Ci\mathcal{C}_{i}. Technically, the captioning models are trained to maximize the joint probability of predicting the ground-truth language tokens, or equivalently minimize the following negative log-likelihood (NLL) loss,

where θ\theta denotes the parameters of the model, and hIi\mathbf{h}_{\mathcal{I}_{i}} denotes the extracted image features of Ii\mathcal{I}_{i}. Previous works like ClipCap fit a powerful text generation model and visual encoding model into this image captioning pipeline. Specifically, strong visual encoding models, such as CLIP , are used to extract the visual features from the input image zi=fCLIP(Ii)\mathbf{z}_{i}=f_{\text{CLIP}}(\mathcal{I}_{i}), then a visual mapping network MV\mathcal{M}_{\text{V}} is trained to map the visual features to ‘prompt vectors’ that adapt to the text generation model, hIi=MV(zi)\mathbf{h}_{\mathcal{I}_{i}}=\mathcal{M}_{\text{V}}(\mathbf{z}_{i}). Finally these prompt vectors hIi\mathbf{h}_{\mathcal{I}_{i}} are fed to a pretrained text generation model, such as GPT , for the captioning task. We adapt this visual captioning pipeline, which uses pretrained feature extractor CLIP and langauge model GPT, for movie AD generation and propose key components that support contextual understanding.

2 Benefiting from Temporal Context

Here, we describe how we extend this single-frame captioning model to include different forms of context, including multiple frames, previous AD text, and subtitles. Compared to image captioning where the annotation describes ‘what is in the image’, movie AD describes the visual happenings in the scene that are relevant to the broader story – often centered around events, characters and the interactions between them. Factors like these cannot be accurately described from a static image alone and therefore a successful automatic AD system must utilize the context of prior events and character interactions.

To tackle these temporal dependencies, we propose to include three components to incorporate the essential contextual information from movies: (i) immediate visual context in the current movie clip (multiple frames), (ii) the previous movie AD, and (iii) the movie subtitles. The architecture of our model is shown in Fig. 2.

Multiple frames (immediate visual context). In contrast to the image captioning method, the visual mapping network MV\mathcal{M}_{\text{V}} takes as input multiple frame features from the current movie clip xi\mathbf{x}_{i} rather than a single image feature, and outputs prompt vectors for the movie clip,

In detail, the mapping network consists of a multi-layer transformer encoder that enables modelling temporal relations among multiple frame features, as shown in Fig 2.

Previous AD text. The sequence of events leading up to the present contain contextual information which are crucial for generating AD of current scene that helps the viewer follow the story. We input this contextual knowledge to our model in the form of the past ADs. Specifically, our model takes the past KK movie ADs {Ti−K,...,Ti−1}\{\mathcal{T}_{i-K},...,\mathcal{T}_{i-1}\} to generate the AD for the current clip. The past movie ADs are a few sentences, which are first concatenated into a single paragraph, then tokenized and converted to a sequence of word embeddings. Inspired by the design of special tokens in language models, we wrap the context AD embeddings with learnable special tokens to indicate the beginning and end of the AD sequence. Formally, the contextual AD embedding is a sequence,

Previous subtitles. Our model also takes the movie subtitles as additional contextual information, which can be sourced either from the official movie metadata or automatically transcribed with an ASR model. The character dialogues, contained with the subtitles, provide complementary information to movie description, including the character names, relationships and emotions. Similar to the context ADs, we concatenate multiple subtitle sentences into a single paragraph and wrap them with learnable special tokens. Practically, since the timing of movie AD does not overlap with the subtitles, we take the most recent LL subtitles within a certain time range as the context,

Due to the weak correlation between the subtitles and the visual elements in the scene, we also experiment with a variant that only encodes the character names occurring in the recent subtitles.

Summary. Overall, the movie AD for the current movie clip Txi\mathcal{T}_{\mathbf{x}_{i}} is generated by conditioning on all the previously described visual and contextual information using a pretrained GPT. The conditional information is fed to GPT as prompt vectors as shown in Fig. 2. The model is trained with NLL loss,

During training, we input the ground-truth past AD. During inference, we experiment with two methods to incorporate the past AD: an oracle setting where the ground-truth past ADs are used in Eq. 2 to generate the current AD, and a recurrent setting where the predicted past ADs are used instead.

3 Pretraining with Partial Data

A major challenge for generating AD is the lack of training data, since the model requires the corresponding visual, textual and contextual data to all be jointly trained. However since our model is modular, components of it can be pretrained with partial data – when a certain type of data is missing, the remaining modules can still be trained. We experiment with partial-data pretraining under two settings: visual-only pretraining and AD-only pretraining.

Visual-only Pretraining. In the absence of contextual data, the visual mapping network can be pretrained with abundant image captioning or (short) video captioning datasets. In this case, the context modules (both contextual AD and subtitles) are deactivated. The training objective of Eq. 2 is turned into L=−log⁡pΘ(Txi∣hxi)\mathcal{L}=-{\log{p_{\Theta}{(\mathcal{T}_{\mathbf{x}_{i}}|\mathbf{h}_{\mathbf{x}_{i}})}}} for visual-only pretraining. Note that the language model is kept frozen here since we find image/video captioning datasets have a clear domain gap with movie AD in both the vision and text modalities.

AD-only Pretraining. Movie AD datasets with corresponding visual information (e.g. frames or frame features) are limited at scale due to potential copyright issues. However, abundant text-only movie ADs are available online as described in Sec. 5. In the absence of visual data, the contextual AD module and the language model can still be pretrained. The training objective in this case becomes L=−log⁡pΘ(Txi∣hAD)\mathcal{L}=-{\log{p_{\Theta}{(\mathcal{T}_{\mathbf{x}_{i}}|\mathbf{h}_{\text{AD}})}}}, which is similar to training a story completion objective by finetuning GPT on text-only movie AD data but with a few additional special tokens. This text-only movie AD pretraining is also related to , which shows a second stage of language model pretraining on in-domain data improves downstream performance.

Denoising MAD Dataset

Our main objective is to generate movie audio descriptions. For this goal, the model is trained on the MAD training set , a dataset of AD caption-video clip pairs from 488 movies. MAD provides the video data in the form of CLIP visual features in order to avoid copyright restrictions. The AD annotations for each movie are automatically collected from AudioVaulthttps://audiovault.net, a large open-source database of audio files containing the full-length original movie track mixed with the AD narrator’s voice. The MAD authors transcribe a subset of this data using ASR, and also have access to the official DVD subtitles. Their automated method then uses text-based speaker separation of the transcribed audio by using subtitles to know when dialogue is present, and assuming all other speech is AD.

This however introduces significant noise because (i) the outdated ASR model results in erroneous transcriptions; and (ii) official DVD subtitles are not exhaustive of all speech in the movie and thus such a method frequently misidentifies character dialogue as AD narration (an example is provided in Fig. 3). Further, obtaining official subtitles from DVDs presents additional challenges when collecting this data at scale.

We propose an improved automated data collection method for AD, requiring only the audio track as input (no DVD subtitles), that tackles both issues by using audio-based speaker separation and an improved ASR model. We then use this method to collect improved annotations for the MAD dataset. Briefly, taking the mixed audio containing both AD narrations and original movie sound track as input, our automated AD collection pipeline contains five stages: (1) speech recognition using WhisperX resulting in punctuated transcriptions with word-level timestamps; (2) sentence tokenization using nltk to provide sentence-level segmentation; (3) speaker diarization to assign speaker labels to each sentence, where the sentence timestamps are used as oracle voice-activity-detection (VAD); (4) labelling the speaker ID of the AD narrator by selecting the cluster with the lowest proportion of first-person pronouns (e.g. ‘I’ and ‘we’); and finally (5) synchronization of the segment timestamps with the visual features by comparing audio. Further details are in the Appendix.

Henceforth we refer to the original MAD annotation as MAD-v1 and our new denoised annotations as MAD-v2. A qualitative comparison is shown in Fig. 5, we find that our MAD-v2 is much more robust and contains less errors and less character dialogue leakage. Both LSMDC and MAD-v1 post-process their annotations by replacing character names in the annotations with ‘someone’ via entity recognition, and release both variants of annotations which we refer to as Named and Unnamed. Similarly, we propose two variants of our denoised annotations: MAD-v2-Named: It contains the raw collected AD narrations without any post-processing on the character names. MAD-v2-Unnamed: Following the character name anonymisation performed in earlier works, we identify character names using a Named Entity Recognition (NER) model and replace them with ‘someone’.

Partial Pretraining with AudioVault Dataset

Paired AD and corresponding visual data are difficult to obtain especially due to movie copyrights, whereas a large number of movie ADs audio tracks are available online for free (e.g. AudioVault). To demonstrate the effect of partial pretraining in Sec. 3.3, we collect a large-scale text-only movie AD dataset from AudioVault. In detail, we source mixed audio files from over 7,000 movies from AudioVault that are not included in MAD-v1, and use a denoising pipeline similar to that described in Sec. 4 to obtain the movie ADs (detailed in Appendix). Additionally we obtain a proxy for the movie subtitles by assuming the ASR from all the non-AD speakers are the characters’ dialogues. To ensure no test-time leakage, we remove all movies present in either LSMDC or MAD from the dataset.

Overall, our AudioVault dataset is an order of magnitude larger than prior AD datasets (see Table 1), from which we provide two sets of data:

AudioVault-AD. The AD narrations from AudioVault and their corresponding timestamps within each movie, totalling 3.3 million AD utterances.

AudioVault-Sub. The subtitles data from AudioVault and their corresponding timestamps within each movie, totalling 8.7 million subtitle utterances.

Experiments

In this section we first outline the experimental details for the AD task, the datasets used for training & testing, the architectural details, and the evaluation metrics (Sec. 6.1). We then report results and discuss the findings, perform ablations on our model, and compare to prior works (Sec. 6.2).

CC3M (Conceptual Caption) is a large image alt-text dataset that contains 3.3M web images. WebVid is a large video-caption dataset that contains 2.5M short stock footage videos. We use them for the partial-data pretraining for visual modules. Additionally, we use our AudioVault-AD to pretrain the textual modules, as described in Sec. 3.3. For the main Movie AD task, we train with original MAD-v1 and our cleaned version MAD-v2, detailed in Sec. 4. Test Datasets. LSMDC contains 118K short video clips with descriptions from 202 movies, of which 182 of them are public. The original MAD-val&test split inherits LSMDC annotations after filtering out 20 lower-quality movies, resulting in 162 movies from all the LSMDC-train/val/test splits. We propose an evaluation split named MAD-eval by further excluding LSMDC train&test movies from these 162 movies, which gives a subset consisting of 10 movies. The reason is twofold: (i) LSMDC-train is commonly used by other works as training data, and (ii) the character names of LSMDC-test are not public. Similarly, we use both MAD-eval-Named and MAD-eval-Unnamed versions. The ‘Unnamed’ version corresponds to the standard LSMDC annotation style – where the characters’ titles and names in the descriptions are replaced by the word ‘someone’; the ‘Named’ version is constructed from the original character names provided by LSMDC. Additionally, subtitles are not provided with MAD-val/test or LSMDC, so we transcribe them from the full-length audio tracks using WhisperX .

1.2 Architecture

For visual features, we use the CLIP ViT-B-32 model , which is a 12-layer transformer encoder that outputs 1×5121\times 512 feature vectors for each input frame. These features are provided by the MAD dataset. For the visual mapping network, we use a 2-layer transformer encoder with 8 attention heads and 512 hidden dimensions, followed by a linear projection layer that projects 512512-d features into 768768-d. We use ten prompt vectors. For the language model, we use GPT-2 , specifically the version from HuggingFace. The GPT-2 model takes as input 768768-d token embeddings, passes through a 12-layer transformer with a causal attention map, and outputs the next token embedding for every input token. We limit the generated number of tokens to 36, since most movie ADs are less than 36 tokens. The GPT-2 is frozen in most of our experiments unless otherwise stated. Each special token (e.g. BAD\texttt{B}_{\text{AD}}) is a learnable 768768-d vector. We take at most 64 past AD tokens and 32 subtitle tokens, and short text samples are padded. Specifically for subtitles, we take the most recent four dialogues within a one-minute time window.

1.3 Training and Inference Details

On the MAD-v1 and MAD-v2 datasets, we use a batch size of 8 sequences, each of which contains 16 consecutive video-AD pairs from a movie. Overall that gives 8×168\times 16 video-AD pairs for every batch. From each video clip, 8 frame features are uniformly sampled. By default, the model is trained for 10 epochs. One epoch means the model has seen all the audio descriptions once. Additional implementation details are in the Appendix.

We use the AdamW optimizer and a cosine-decay learning rate schedule with a linear warm-up. The starting learning rate is 10−410^{-4} and is decayed to . For each experiment, we use a single Nvidia A-40 for training. For text generation, greedy search and beam search are commonly used sampling methods. We stop the text generation when a full stop mark is predicted, otherwise we limit the sequence length to 67 tokens. We use beam search with a beam size of 55 and mainly report results by the top-1 beam-searched outputs, since beam search performs slightly better than greedy search on multiple scenarios. Note that under the ‘recurrent’ setting, we feed the past greedy-searched text outputs to the model to generate the current AD, which we find gives more stable results.

1.4 Evaluation Metrics

To evaluate the quality of text compared with the ground-truth, we use classic metrics including ROUGE-L (R-L), CIDEr (C) and SPICE (S). We also report BertScore (BertS), which evaluates word matching between a candidate sentence and reference sentence with pretrained BERT embeddings. A higher value indicates better text generation compared with the ground-truth.

2 Experiments on Movie Audio Descriptions

In Table 2 we show that visual context from multiple frames brings a clear gain for the AD task (C 6.7 vs 4.0). AD context provides a consistent performance improvement under both oracle (C 17.8 vs 6.7) and recurrent settings (C 12.6 vs 6.7). Note that we find feeding AD context as text tokens works better than training a textual feature mapping network, we conjecture the ADs in their original text form carry the most key information like the names and places. However, subtitle context provides no gain for our model (C 13.3 vs 14.3) under the recurrent setting, which we attribute to the very weak correspondence between the visual elements in the scene and the character dialogue. When the subtitles are filtered and contain only character names (denoted as ‘SubN’), they provide a slight performance gain (C 14.2 vs 13.3). Since the subtitles used are without speaker identities, the model may struggle to know which character in the frame spoke each subtitle. Overcoming these challenges will be considered in future work.

Table 3 demonstrates the benefit of our MAD v2 annotations over v1, confirming the qualitative findings. Training the AD model with context on v2 outperforms training on v1 under all settings (both named and unnamed) by a significant margin. Since the v2 annotations are fewer in number than MAD-v1, this suggests they are indeed less noisy and result in AD captioning models with improved performance.

In Table 2, we find that visual-only pretraining on open-domain vision-text data provides clear gains (CIDEr 8.4 vs 6.7 for CC3M, and 10.0 vs 6.7 for WebVid). But considering the size of visual samples, the improvement is not data-efficient. We attribute this to the large domain gap between movie AD and classical visual caption annotations like CC3M or WebVid2M. The text-only pretraining of our model also improves performance. For the recurrent AD context model, AudioVault-AD pretraining increases CIDEr from 12.6 to 14.1, which indicates the great importance of adapting to the text style and context. The combination of the visual module after visual-only pretraining (WebVid) and the textual modules after text-only pretraining (AV-AD) gives a further performance gain (C 21.9 vs 19.0 for the oracle setting, and 14.3 vs 14.1 for recurrent).

In Figure 4 we show the effect of varying the number of context ADs given to the model. Longer AD context improves performance almost consistently across all settings, but it brings extra computational cost due to the quadratic complexity of the attention operation in GPT-2. Note that we experiment with at most 6 contextual AD sentences, which is equivalent to about 70-word embeddings in Eq. 1. The trend for the recurrent setting flattens when the context ADs are longer than 3 sentences, which is probably due to the limited power of processing long context for the GPT2 model.

2.1 Qualitative Results

Fig. 5 shows qualitative examples of our model. Under the oracle setting, the model can use the character identities easily from the past ground-truth AD (e.g. “master-at-arms”). Whereas under the recurrent setting, the model can only learn names from the subtitles but names appear very sparsely in subtitles, therefore the model mostly predicts pronouns (e.g. “he”, “they”) but still gets the actions (“looks”) or objects (“necklace”) correct.

3 Comparison with Other Works

In Table 4, we compare our method with previous visual captioning methods. Note that since the MAD dataset only releases the CLIP visual features, rather than the movie frames, our comparison is limited to methods that build on frozen CLIP features. We show a clear performance improvement compared to ClipCap and CapDec , for the latter the language model is also adapted to the movie AD domain by text-only pretraining. The results highlight the importance of context for movie AD.

In Table 5, we adapt our method to the Multi-Sentence Description task on LSMDC, in which the model takes five consecutive clips and generates five corresponding descriptions. Since the task is performed on the unnamed annotations, we finetune our best model in Table 4 with varying 0-4 context ADs as input on MAD-v2-Unnamed dataset and test with the recurrent setting. To make minimal changes, our model still takes a single clip feature at each step, whereas previous methods take all five clips together for movie description. Despite this disadvantage, we obtain competitive results on this task even without using the manually-cleaned LSMDC training set (C 16.7 vs 15.4), effectively zero-shot. The performance of the model can be further improved by additionally training on LSMDC data.

Conclusion and Future Work

This paper focuses on the automatic generation of movie AD for a given time interval, and has made significant progress. We propose an AutoAD pipeline that incorporates contextual information. Additionally, we demonstrate the effectiveness of partial-data pretraining, a technique that could be widely applicable when full data is difficult to obtain. Further, we clean up the previous MAD dataset and collect a new text-only movie AD dataset as a pretraining resource. However, a clear limitation of this AutoAD pipeline is character naming – referencing who is doing what, a necessary ingredient for story-coherent movie AD. Additionally, future work could tackle the problem of when to generate AD, instead of relying on the annotated AD timestamps.

Acknowledgements. We thank Mattia Soldan for helping with the MAD dataset, Anna Rohrbach for the LSMDC dataset, and the AudioVault team for their priceless contribution to the visually impaired. This research is funded by EPSRC PG VisualAI EP/T028572/1, a Google-Deepmind Scholarship, and ANR-21-CE23-0003-01 CorVis.

References

Appendix A AD Collection Pipeline Additional Details

Collecting movie AD has two main challenges. First, in the audio files (e.g. from AudioVault) the movie AD is fused with the original movie audio, i.e. on the same audio track. The pipeline needs to identify the AD speaker among the movie characters accurately. Second, for the same movie, the audio files from AudioVault is usually not synchronised with the movie from which the MAD visual features were extracted, mainly due to the varied durations of intro and outro of different movie source. Since we rely on the MAD visual features, the synchronisation is an essential step.

The automated data collection pipeline is briefly introduced in Sect. 4 of the main paper. A schematic is shown in Fig. A.1, in detail:

We transcribe the mixed audio file using WhisperX which provides accurate punctuated transcriptions with word-level timestamps.

The transcript is tokenized into sentences using the nltk python toolbox , resulting in transcription sentences and their corresponding temporal segments (inferred the start and end time of the first and last word in the sentence respectively).

Each sentence segment is assigned a single speaker identity (e.g. SPEAKER_00, SPEAKER_01, etc.) by performing speaker diarization on the mixed audio, whereby each sentence timestamp is provided as oracle voice activity detection. Specifically, we use SpeechBrain ECAPA-TDNN voice embeddings trained on VoxCeleb and Agglomerative Clustering with a threshold of 0.95.

To automatically identify the cluster associated with the AD speaker, we exploit the third-person nature of AD narrations and select the cluster with the lowest proportional occurrence of first- & second-person pronouns, e.g. “I” and “you” with 95 or more speaker segments.

To synchronise the segment timestamps with the original audio track from which the MAD visual features were extracted, we follow and calculate the time delay τ\tau between the original movie audio files and the mixed audio files via FFT cross-correlation. The timestamps of the identified AD segments are shifted according τ\tau in order to synchronise them to the visual features and subtitles collected in MAD.

A.2 AD Collection Pipeline for AudioVault

The collection pipeline for AudioVault is introduced in Sect. 5 of the main paper, we provide more details here. To collect text-only AD annotations from AudioVault, the final synchronisation step is unnecessary. Therefore, we follow steps 1-4 of the MAD denoising pipeline as described above, which takes as input the mixed audio tracks and outputs the ASR with timestamps from the possible AD speaker.

The large-scale collection from AudioVault audio files is noisy, e.g. some ADs are of lower-quality or are sourced from short movies. Therefore, we apply a stricter filtering step that removes movies containing fewer than 100 AD narrations or a word frequency of first- & second-person pronouns larger than 5%.

A.3 Comparison with MAD-v1

The key advantages of our pipeline are three-fold: (1) it relies on audio-based speaker separation to identify the AD speaker among the movie characters, whereas the pipeline in the original MAD work relies on text-based speaker separation by using the timestamps from the DVD subtitles and assumes any ASR transcription outside of these timestamps is AD. The error is propagated because the official subtitles are non-exhaustive (some dialogue is missed by the official subtitles). (2) It requires only the mixed audio as input, whereas MAD must also source the official DVD subtitles and align them – presenting additional scaling costs and challenges. (3) It uses an advanced ASR model Whisper which gives much more accurate transcriptions than previous methods, especially for punctuation and the spelling of names and other identities.

Appendix B Qualitative Examples of MAD-v2 vs MAD-v1.

More qualitative examples of MAD-v2 and MAD-v1 are shown in Fig. A.2 and A.3. It is clear that our pipeline produces more accurate AD compared to the original MAD-v1, particularly in the spelling of names and the exclusion of dialogue.

Appendix C Quantitative Comparison between MAD-v2 vs MAD-v1 on Grounding

We re-purpose the CLIP zero-shot video-language grounding (VLG) performance from as an indicator of dataset quality. In detail, for both MAD-v2 and MAD-v1, we randomly choose a set of 5 movies from the training split, and compute the VLG performance with frozen CLIP visual and textual encoders. The AD textual quality and timestamps are the only factors that differ in this comparison. We use the MAD training split because we did not modify the val/test splits, which are from LSMDC annotations. The code to compute VLG performance is from https://github.com/Soldelli/MAD. The result in Table A.1 shows MAD-v2 annotations also benefit the VLG task.

Appendix D Additional Implementation Details

Number of frames per movie clip NN: We choose N=8N=8. Most AD annotations have a time duration of 1-3 seconds, equivalent to 5-15 frames (features) under 5 FPS – the sampling rate provided by the MAD dataset. Therefore N=8N=8 is a reasonable choice.

Number of AD sentences as context KK: We experiment K∈{1,2,3,6}K\in\{1,2,3,6\} in Figure 4.

Number of subtitles LL: As described in Sect. 6.1.2, for simplicity we take the most recent 4 dialogues within a 1-minute time window. Note that the time distribution of subtitles varies a lot – the most recent 4 dialogues could span just a few seconds or up to minutes before the current AD timestamp.

We use the pycocoeval package from https://github.com/tylin/coco-caption to compute the ROUGE-L, CIDEr, SPICE and METEOR. The package post-processes both the predicted text and ground-truth text internally to remove the punctuation and make them lowercase. To compute the BertScore, we use the package from https://github.com/Tiiiger/bert_score. Note that before computing the BertScore, both the predicted text and the ground-truth text are converted to lowercase without any punctuation, as these are factors that the BertScore is sensitive to.

We investigate an alternative vision & language fusion mechanism whereby the context AD sentence prompts are fed as language features rather than raw text tokens. Empirically, we observe that raw text inputs outperform language features (e.g. 12.6 CIDEr in Table 2 vs. about 8.0 CIDEr when feeding language features).

Appendix E Additional Qualitative Examples

More qualitative examples are shown in Fig. A.4. It shows that the AutoAD model gives reasonable descriptions for the movie domain, like the actions (swim, dance), and face expression (eyes widen). Note that under the oracle setting, the model is capable of learning character names (sample a, c, d, f) mainly due to the extra information from the ground-truth context. The model is still limited in its ability to identify characters accurately, e.g. in sample f, the movie shows Nick and Daisy are dancing. Whereas the oracle prediction describes that Nick and Gatsby are dancing, and the recurrent model simply predicts that a man and woman are dancing. Also the pronouns often appear in the recurrent prediction, such as the word ‘his’ in sample a and b, which shows the model learns the bias of pronouns but cannot recognize characters correctly.

Appendix F Dataset Splits

To clarify the dataset split of the movies in MAD and LSMDC, we list the movie IDs of each split used (and not used) in this paper. The splits can also be found on the website https://www.robots.ox.ac.uk/~vgg/research/autoad/.

It consists of 488 movies and all of them are used for training. We provide the cleaner ADs for these movies using the automated pipeline described above. This set is the same as the training set of movies proposed in MAD . The movie IDs are:

It consists of 10 movies, which are obtained by set(MAD val/test)∩set(LSMDC val)set(\text{MAD val/test})\cap set(\text{LSMDC val}), that excluding LSMDC train movies for the ease of future comparison. The annotations are inherited from the LSMDC dataset, and we use both the named and unnamed version of it, where the named version can be downloaded from the LSMDC websitehttps://sites.google.com/site/describingmovies/download?authuser=0. The movie IDs are: [1005_Signs, 1026_Legion, 1027_Les_Miserables, 1051_Harry_Potter_and_the_goblet_of_fire, 3009_BATTLE_LOS_ANGELES, 3015_CHARLIE_ST_CLOUD, 3031_HANSEL_GRETEL_WITCH_HUNTERS, 3032_HOW_DO_YOU_KNOW, 3034_IDES_OF_MARCH, 3074_THE_ROOMMATE]

151 movies from MAD val/test are not used in either training or testing in our paper. They are the intersection set(MAD val/test)∩set(LSMDC train/test)set(\text{MAD val/test})\cap set(\text{LSMDC train/test}). The movie IDs are: [0001_American_Beauty, 0002_As_Good_As_It_Gets, 0003_CASABLANCA, 0004_Charade, 0005_Chinatown, 0006_Clerks, 0007_DIE_NACHT_DES_JAEGERS, 0008_Fargo, 0009_Forrest_Gump, 0010_Frau_Ohne_Gewissen, 0011_Gandhi, 0012_Get_Shorty, 0013_Halloween, 0014_Ist_das_Leben_nicht_schoen, 0016_O_Brother_Where_Art_Thou, 0017_Pianist, 0019_Pulp_Fiction, 0020_Raising_Arizona, 0021_Rear_Window, 0022_Reservoir_Dogs, 0023_THE_BUTTERFLY_EFFECT, 0026_The_Big_Fish, 0027_The_Big_Lebowski, 0028_The_Crying_Game, 0029_The_Graduate, 0030_The_Hustler, 0031_The_Lost_Weekend, 0032_The_Princess_Bride, 0033_Amadeus, 0038_Psycho, 0041_The_Sixth_Sense, 0043_Thelma_and_Luise, 0046_Chasing_Amy, 0049_Hannah_and_her_sisters, 0050_Indiana_Jones_and_the_last_crusade, 0051_Men_in_black, 0053_Rendezvous_mit_Joe_Black, 1001_Flight, 1002_Harry_Potter_and_the_Half-Blood_Prince, 1003_How_to_Lose_Friends_and_Alienate_People, 1004_Juno, 1006_Slumdog_Millionaire, 1007_Spider-Man1, 1008_Spider-Man2, 1009_Spider-Man3, 1010_TITANIC, 1011_The_Help, 1012_Unbreakable, 1014_2012, 1015_27_Dresses, 1017_Bad_Santa, 1018_Body_Of_Lies, 1019_Confessions_Of_A_Shopaholic, 1020_Crazy_Stupid_Love, 1028_No_Reservations, 1031_Quantum_of_Solace, 1033_Sherlock_Holmes_A_Game_of_Shadows, 1034_Super_8, 1035_The_Adjustment_Bureau, 1037_The_Curious_Case_Of_Benjamin_Button, 1038_The_Great_Gatsby, 1039_The_Queen, 1040_The_Ugly_Truth, 1042_Up_In_The_Air, 1043_Vantage_Point, 1045_An_education, 1046_Australia, 1047_Defiance, 1048_Gran_Torino, 1050_Harry_Potter_and_the_deathly_hallows_Disk_One, 1052_Harry_Potter_and_the_order_of_phoenix, 1054_Harry_Potter_and_the_prisoner_of_azkaban, 1055_Marley_and_me, 1057_Seven_pounds, 1058_The_Damned_united, 1059_The_devil_wears_prada, 1060_Yes_man, 1061_Harry_Potter_and_the_deathly_hallows_Disk_Two, 1062_Day_the_Earth_stood_still, 3001_21_JUMP_STREET, 3002_30_MINUTES_OR_LESS, 3003_40_YEAR_OLD_VIRGIN, 3004_500_DAYS_OF_SUMMER, 3005_ABRAHAM_LINCOLN_VAMPIRE_HUNTER, 3007_A_THOUSAND_WORDS, 3008_BAD_TEACHER, 3012_BRUNO, 3013_BURLESQUE, 3014_CAPTAIN_AMERICA, 3016_CHASING_MAVERICKS, 3017_CHRONICLE, 3018_CINDERELLA_MAN, 3020_DEAR_JOHN, 3022_DINNER_FOR_SCHMUCKS, 3023_DISTRICT_9, 3024_EASY_A, 3025_FLIGHT, 3026_FRIENDS_WITH_BENEFITS, 3028_GHOST_RIDER_SPIRIT_OF_VENGEANCE, 3030_GROWN_UPS, 3033_HUGO, 3035_INSIDE_MAN, 3036_IN_TIME, 3037_IRON_MAN2, 3038_ITS_COMPLICATED, 3039_JACK_AND_JILL, 3040_JULIE_AND_JULIA, 3041_JUST_GO_WITH_IT, 3042_KARATE_KID, 3043_KATY_PERRY_PART_OF_ME, 3045_LAND_OF_THE_LOST, 3046_LARRY_CROWNE, 3047_LIFE_OF_PI, 3048_LITTLE_FOCKERS, 3049_MORNING_GLORY, 3050_MR_POPPERS_PENGUINS, 3051_NANNY_MCPHEE_RETURNS, 3052_NO_STRINGS_ATTACHED, 3053_PARENTAL_GUIDANCE, 3054_PERCY_JACKSON_LIGHTENING_THIEF, 3055_PROMETHEUS, 3056_PUBLIC_ENEMIES, 3058_RUBY_SPARKS, 3060_SANCTUM, 3061_SNOW_FLOWER, 3062_SORCERERS_APPRENTICE, 3063_SOUL_SURFER, 3066_THE_ADVENTURES_OF_TINTIN, 3067_THE_ART_OF_GETTING_BY, 3069_THE_BOUNTY_HUNTER, 3070_THE_CALL, 3071_THE_DESCENDANTS, 3072_THE_GIRL_WITH_THE_DRAGON_TATTOO, 3073_THE_GUILT_TRIP, 3075_THE_SITTER, 3076_THE_SOCIAL_NETWORK, 3077_THE_VOW, 3078_THE_WATCH, 3079_THINK_LIKE_A_MAN, 3081_THOR, 3082_TITANIC1, 3083_TITANIC2, 3084_TOOTH_FAIRY, 3085_TRUE_GRIT, 3086_UGLY_TRUTH, 3087_WE_BOUGHT_A_ZOO, 3088_WHATS_YOUR_NUMBER, 3089_XMEN_FIRST_CLASS, 3090_YOUNG_ADULT, 3091_ZOMBIELAND, 3092_ZOOKEEPER]