Streaming Dense Video Captioning

Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, Cordelia Schmid

Introduction

Video is ubiquitous in modern society, quickly becoming one of the most prevalent media formats for transmitting information. The majority of computer vision models designed for video understanding only process a handful of frames, mostly covering only a few seconds , and are typically limited to classifying these short segments into a fixed number of concepts. In order to achieve a comprehensive, fine-grained video understanding, we study the task of dense video captioning – jointly localizing events temporally in video and generating captions for them. Ideal models for this goal should be able to handle both long input sequences – to reason over long, untrimmed videos – and also to handle long output sequences in text space, to describe in detail all the events within the video.

Prior work on dense video captioning does not handle either long inputs or long outputs. Given a video of any length, state-of-the-art models either sample very few frames (e.g., 6 frames ) with large strides (i.e., temporal downsampling), or keep one feature per-frame for all the frames (i.e., spatial downsampling). With long textual output, current models solely rely on auto-regressive decoding to generate multiple sentences at the end.

In this work, we design a streaming model for dense video captioning as shown in Fig. 1. Our streaming model does not require access to all input frames concurrently in order to process the video thanks to a memory mechanism. Moreover, our model can produce outputs causally without processing the entire input sequence, thanks to a new streaming decoding algorithm. Streaming models such as ours are inherently suited to processing long videos – as they ingest frames one at a time. Moreover, as the output is streamed, intermediate predictions are produced before processing the full video. This property means that streaming models can in theory be applied to process live video streams, as required for applications such as video conferencing, security and continuous monitoring among others.

In order to develop our streaming model, we first propose a novel memory mechanism, that takes frames once at a time. The memory model is based on K-means clustering, and uses a fixed number of cluster-center tokens to represent the video at each timestamp. We show that this method is simple and effective, and can process variable numbers of frames, with a fixed computational budget at decoding.

We also develop a streaming decoding algorithm, and train our network such that given a “decoding point” (Fig. 2) at a particular timestamp, it predicts all event captions that ended before it given the memory features at that timestamp. Our network is thus trained to make predictions at any timestamp of the video, and not just at the end of the video as in conventional, non-streaming models. Furthermore, we provide predictions from earlier decoding points as contexts for later ones. This context avoids predicting duplicated events, and can be used as an “explicit” memory in natural language summarising the earlier video. Our streaming output is also motivated by the fact that as the video length grows, our memory will inevitably lose information over time as its size is bounded. We avert this issue by making predictions before we have processed the entire video, and still keep early information via language context.

We evaluate our approach on three popular dense video captioning datasets, ActivityNet , YouCook2 and ViTT . Our results show that our streaming model significantly improves over the state-of-the-art, which inevitably uses fewer frames or fewer features, by up to 11.0\mathbf{11.0} CIDEr points. We show that our method generalizes across both GIT and Vid2Seq architectures. Finally, our proposed memory can also be applied to paragraph captioning, improving baselines by 11-55 CIDEr points.

Related Work

Dense video captioning. Dense video captioning requires captioning events, and localizing them temporally. Traditionally, prior work used a two-stage approach, first localizing events in video, and then subsequently captioning them . More recent end-to-end approaches include PDVC which infers event captions and timestamps using a DETR-like model. Vid2Seq augments the vocabulary of a language model with timestamp tokens, allowing them to generate concatenated event captions in the same manner as a regular captioning model. We also use the output formulation of , as it integrates well with foundation vision-language models . A related, but different problem, is that of audio description in movies , which requires generating captions for the visually impaired that must be complementary to speech, and often uses auxiliary, non-causal models to recognize characters or their speech .

As far as we are aware, all prior dense captioning models are not causal, as they encode the entire video at once. Moreover, to process long videos, they typically use visual features that are downsampled heavily (by selecting a few frames , or spatially pooling features per frame ). In contrast, we process videos in a streaming manner, processing long input sequences one frame at a time with a memory module, and streaming output sentences with a novel decoding algorithm.

Models for long videos. A common way of processing longer videos is to use a memory mechanism to provide a compact representation of past events. With transformers, memory can easily implemented by using tokens from past observations as inputs for the present time step . Examples in vision include which pre-extract features offline and retrieve them during inference, and are thus not causal. MemViT uses token activations from previous time steps as inputs to the current time step. However, this means that the sequence length grows over time and so it cannot handle arbitrarily long videos.

An alternate view of memory is to see it as a method of compressing previously observed tokens into a smaller, fixed-size set to be used at future time steps. Token Turing Machines summarize past and current observations using the token summarization module of . MovieChat follows a similar idea, but uses a variant of Token Merging instead to perform the summarization. TeSTra uses an exponential moving average to integrate video features instead. The advantage of such approaches is that the memory bank has a fixed size, and therefore the computational cost remains bounded regardless of the length of the video. Our memory model has this same desirable property. However, our memory is based on clustering, using the centers from a K-means-like algorithm to summarize tokens from each time step, and we show experimentally that this outperforms other alternatives.

Causal models in video. Our streaming model is causal, meaning that its output only depends on current and past frames, without access to future frames. Although we are not aware of prior causal models for dense video captioning, there are causal models in many other vision domains. Online action detection aims to predict action labels for videos in real-time without access to future frames. Analogously, online temporal action localization models also predict start- and end-times after an action is observed. Most models for object tracking and video object/instance segmentation are causal too.

A common theme in the above tasks is that the model must make a prediction at each frame of the video. In contrast, we focus on dense video captioning , which is challenging as the output captions do not have a one-to-one correspondence with the video frames. We address this problem by proposing a streaming decoding algorithm.

Streaming Dense Video Captioning

Captioning models broadly consist of a vision encoder followed by a text decoder. We outline these approaches, and show how they can be extended to dense captioning next.

Text decoder. Given the visual features, f\mathbf{f}, and optional textual prefix tokens, p\mathbf{p}, the text decoder, D\mathcal{D} generates a sequence of word tokens, c\mathbf{c} from them. We use an autoregressive decoder that generates the next word token, wiw_{i}, conditioned on previous words, w1:i−1\mathbf{w}_{1:i-1}, and prefix if provided as wi=D(f,p,w1:i−1)w_{i}=\mathcal{D}(\mathbf{f},\mathbf{p},\mathbf{w}_{1:i-1}). Note that prefix tokens are typically not used in captioning tasks, but are used in question-answering (QA) to encode the input question. Concretely, the text decoder, D\mathcal{D}, is a sequence of transformer layers operating on a concatenation of visual features f\mathbf{f} and word embeddings of the prefix . This architecture is shown to be effective in both captioning and QA tasks across image and video .

Dense video captioning with timestamps. Combining the above visual encoder and text decoder gives a basic architecture for video captioning. To extend it for captioning multiple events with starting and ending timestamps, Vid2Seq introduced two main modifications: First, it augments the vocabulary, V′V^{\prime}, of the captioning model with time tokens, wsw^{s} and wew^{e}, which represent the starting and ending times, respectively. A single event is therefore represented as c′=[ws,we,w1,⋯ ,wn]\mathbf{c}^{\prime}=[w^{s},w^{e},w_{1},\cdots,w_{n}], and ∣V′∣=∣V∣+∣T∣|V^{\prime}|=|V|+|T| where ∣V∣≤ws<we≤∣V′∣|V|\leq w^{s}<w^{e}\leq|V^{\prime}|, and ∣T∣|T| is the number of time tokens. Second, Vid2Seq concatenates all timed captions into a single long caption that is ordered by starting time: C=[c1′,c2′,⋯ ,cne′]\mathbf{C}=[\mathbf{c}^{\prime}_{1},\mathbf{c}^{\prime}_{2},\cdots,\mathbf{c}^{\prime}_{n_{e}}] where nen_{e} is the number of events. Therefore, dense video captioning can be formulated as standard video captioning with target C\mathbf{C}.

Despite its effectiveness, the (generalized) Vid2Seq architecture has a number of key limitations: First, it forwards visual features from the whole video, f\mathbf{f} through the decoder, meaning that it does not scale effectively to longer videos and more tokens. In addition, as Vid2Seq predicts all event captions at once, after processing the whole video, it struggles with predicting long, detailed captions. To address these issues, we introduce streaming dense video captioning models, where we process the inputs once at a time using a memory module to bound computational costs, and stream the outputs such that we can make predictions before processing the whole video.

2 Streaming inputs using memory

Next, we update the memory at each timestamp for each incoming frame ft\mathbf{f}_{t}. Our intuition is to keep as much diverse information in the original video as possible, while not increasing the storage budget (i.e. by keeping a constant memory size KK). We thus propose a K-means-like clustering algorithm, to use the feature cluster centers as the approximate video features. To avoid the cluster centers biasing quickly to incoming features, we keep track of the number of merged tokens in each cluster center. We use this as a momentum weight, so that cluster centers that are merged from more tokens change slower. The detailed algorithm diagram is provided in Alg. 1, and illustrated in Fig. 3.

The K-means algorithm is not differentiable with respect to the assignment of data points to cluster centers, δ\bm{\delta} (Line 1 of Alg.1). However, the inputs and outputs of our memory module are the updated cluster centers, Mt\mathbf{M}_{t}, which is a linear mapping of the input X=[Mt−1,ft]\mathbf{X}=[\mathbf{M}_{t-1},\mathbf{f}_{t}], as Mt=AX\mathbf{M}_{t}=\textbf{A}\mathbf{X}, where A is a weight matrix computed from X. Therefore, even though we cannot compute the gradient of A\mathbf{A} with respect to X\mathbf{X}, we can compute the gradient of Mt\mathbf{M}_{t} with respect to the input X\mathbf{X}, and thus to the input visual feature f\mathbf{f}. As a result, we can use our memory module in any part of a neural network, and learn parameters in preceding layers.

3 Streaming outputs with decoding points

The memory module from Sec. 3.2 enables us to efficiently ingest long input videos. However, it is still desirable for our model’s text decoder to predict outputs before it has processed the entire input sequence: Streaming the output substantially decreases the latency of the model, as we do not have to wait for the model to process the entire input sequence to make predictions. This is particularly relevant for processing, for example, live video streams. Furthermore, streaming the output can in fact increase our model’s accuracy: As we have a memory with a fixed size, KK, from which we decode outputs, we will inevitably lose information over time. We can therefore avert this issue by making predictions before we have processed the entire video.

As shown in Fig. 4, we define “decoding points”, did_{i}, as intermediate timestamps where we decode event captions given the features in our memory, Mdi\mathbf{M}_{d_{i}}. We train our model such that at each decoding point, did_{i}, the model predicts all event captions that finished before it. More specifically,

where Yi\mathcal{Y}_{i} is the set of all event captions corresponding to the ithi^{\text{th}} decoding point did_{i}, and wjsw^{s}_{j}, wjew^{e}_{j} are the starting and ending time of the jthj^{\text{th}} event.

As decoding points are applied sequentially, later decoding points should have access to the predictions of earlier decoding points, and should not repeat them again. Therefore, from the second decoding point onwards, we concatenate the outputs of previous decoding points as the prefix to the text decoder, as shown in Fig. 4. Moreover, during training, we perform further data augmentation by randomly removing some of the previous event captions from the prefix, and adding them to the target instead, to increase robustness to potential errors in earlier predictions. We therefore denote our prefixes and captioning targets during training as

where j<∣Yi∣j<|\mathcal{Y}_{i}| is a randomly chosen splitting point to partition the target and context. During inference, we use the models actual predictions, pi=[y^1,y^2,…,y^i−1]\mathbf{p}_{i}=[\hat{\mathbf{y}}_{1},\hat{\mathbf{y}}_{2},\ldots,\hat{\mathbf{y}}_{i-1}], instead.

In practice, we uniformly sample decoding points during both training and inference, with a stride of SS frames. Since the model is trained to predict all event captions before the decoding point, it means that the exact location at inference time does not need to match the event boundaries closely. The number of decoding points can be also different between training and inference, as we will ablate. This method is both simple and scalable, as we show experimentally in the next section.

Experimental Evaluation

We evaluate our model on the three most popular dense video captioning datasets: ActivityNet Captions , YouCook2, and ViTT .

ActivityNet Captions contains 8,649 training videos and 4,267 validation videos (considering videos that are still online on YouTube). The videos are untrimmed, each containing multiple action events, with an average video length of 2 minutes and an average of 3.7 events. The videos are selected from the ActivityNet , which covers a wide range of visual domains. Each video is carefully labeled by human annotators. We use this dataset for both dense captioning (predicting a sequence of captions with their start- and end-times) and paragraph captioning (predicting the aforementioned captions as a single paragraph without timestamps).

YouCook2 contains 1,333 training videos and 456 validation videos. The videos are on average 5.35.3 minutes long, with an average of 7.87.8 events. This dataset focuses on cooking scenes, and most videos contain one main actor cooking while describing the recipe. The videos are manually annotated, aided by speech transcriptions obtained through Automatic Speech Recognition (ASR). As expected due to the annotation process, using ASR as an additional modality can be used to improve model predictions . However, our work focuses on using visual-only inputs, to study our proposed streaming model. We also perform dense and paragraph captioning on this dataset.

ViTT contains 4,608 training videos and 2,301 testing videos, which are on average 4.74.7 minutes long and contain 77 events per video. It is collected in a similar way as YouCook2, but focuses on more diverse domains. Again, we do not use the ASR inputs in our experiments.

Evaluation Metrics For all datasets, we report standard dense captioning evaluation metrics following Vid2Seq and PDVC . The primary metric is CIDEr which has been adapted to jointly evaluate captioning and localisation. In particular, we average the CIDEr score for positive ground-truth-prediction pairs with a temporal IoU above the thresholds, {0.3,0.5,0.7,0.9}\{0.3,0.5,0.7,0.9\}. Similarly, we report METEOR averaged over multiple IoU thresholds. Finally, we report the recently proposed SODAc , which considers all event captions from the same video, and is therefore a more holistic measure.

1.2 Implementation details

We implement our streaming model on two video captioning architectures, GIT and Vid2Seq . Both use a ViT-L initialized from CLIP as image encoder F\mathcal{F}.

GIT concatenates all tokens from all input frames, f\mathbf{f}, and feeds them to a 6-layer transformer language decoder. We apply our streaming input module before the language decoder, i.e., rather than concatenating frame features, we use our memory features instead. The original GIT paper pretrains the language decoder on in-house data . As we do not have access to the original data or weights, we pretrain our GIT model on the WebLI dataset .

Vid2Seq pools all frame tokens spatially into 1 token per frame, and then applies a 12-layer temporal transformer to obtain the final visual features. The language decoder is a 12-layer T5-Base decoder . Vid2Seq pretrains the temporal transformer and the T5 decoder on narrated videos from the YT-Temporal dataset . We use their public, pretrained weights as our initialization. We apply our streaming input module before the temporal transformer, to ensure the intermediate features remain causal.

When finetuning our model for each dataset, we freeze the frame encoder following prior works . We provide detailed training hyperparameters in the supplementary. For each architecture, the training hyperparameters are shared across the 3 datasets.

2 Analysis of Streaming Modules

We first analyze each of the main components of our model: streaming the input with our memory module (Sec. 3.2), and then streaming outputs using decoding points (Sec. 3.3). Unless otherwise stated, we use the GIT backbone for our experiments. Due to randomness in training and evaluation, for all ablation experiments (Tab. 1-Tab. 3), we repeat the same run 3 times and report the averaged metrics. When compared to other methods (Tab. 4, Tab. 5), we report results for the best model.

As the GIT architecture concatenates visual features, f\mathbf{f}, from multiple frames before passing it to the language decoder, it is limited by the number of frames that it can process. Our clustering-based memory module allows us to process more frames, and we consider the following approaches to evaluate it:

No memory: We simply concatenate all visual tokens from all frames as done in GIT . Due to memory constraints, we can only feed up to 16 frames to the decoder.

Spatial- or temporal-pooling: We pool the visual features, f, along either the spatial or temporal dimensions to reduce the number of tokens fed to the language decoder.

EMA: We use an exponential moving average of frame features, ft\mathbf{f}_{t}, at each time step, following TeSTra . We sweep the decay rate from {0.9,0.99,0.999}\{0.9,0.99,0.999\}, finding 0.90.9 to perform best.

MovieChat . Finally, the recent MovieChat paper maintains a memory of KK tokens. For each incoming frame, it sequentially processes each token, and merges the two most similar tokens in the memory bank such that the size remains fixed at KK. We implement this method using the author’s public code.

Tab. 1 compares the results of the different memory modules. For T ⁣ ⁣= ⁣ ⁣16T\!\!=\!\!16, where we can feed all the tokens from the vision backbone, f\mathbf{f}, into the decoder, “no memory” and our method both performs the best. We expected “no memory” to perform well, as it uses the most tokens, T ⁣ ⁣× ⁣ ⁣NfT\!\!\times\!\!N_{f}. However, our clustering-based method achieves the same performance, which suggests that it is able to effectively capture the relevant information in the video with 8×8\times fewer tokens, and is causal too. It is not possible to use “no memory” for T ⁣ ⁣> ⁣ ⁣16T\!\!>\!\!16 due to its computational cost.

With more frames, naïvely pooling along the spatial-, or temporal-dimensions actually performs worse. This is likely because we are averaging out information over longer temporal durations, and thus losing the details required for more detailed localization or captioning. Our method and MovieChat on the other hand, are able to leverage more frames to improve performance, as they keep diverse features within the memory. Finally, our clustering method outperforms MovieChat and other memory modules for all numbers of frames that we considered, which is why we use it for all future experiments.

We also ablate the hyperparameters of our memory module in Tab. 2. Following this experiment, we set K ⁣= ⁣257 ⁣× ⁣2K\!=\!257\!\times\!2, or the number of tokens in two frames, and use 2 iterations of K-means, as it achieves the best performance. Our proposed momentum term (Sec. 3.2), which prevents cluster centers from becoming too biased towards incoming frame tokens, also improves CIDEr from 29.7 to 30.6.

2.2 Streaming outputs

We now analyze the effect of our streaming decoding method (Sec. 3.3), which enables us to predict event captions from intermediate features in our memory, and utilize previous prediction as the prefix to our text decoder. Table 3 analyses the effects of the number of decoding points during training, the impact of the prefix, and how we select decoding points during inference.

Table 6(a) shows that we achieve significant improvements by increasing the number of decoding points during training, improving the CIDEr score by 1010 points, or 33% relative, compared to only making predictions at the end, as in conventional models. Streaming the output with decoding points can provide performance benefits for multiple reasons: First, as we have multiple decoding points, the required output caption is shorter at each decoding point, making the captioning task easier. Second, the visual features from our memory, M\mathbf{M}, may be aligned better with the target text, since M\mathbf{M} does not represent the visual features from the entire video as in baseline approaches, but only the features up to the decoding point. Finally, training with decoding points provides a stronger training signal to the model, as we provide training supervision at each decoding point. Moreover, it should aid generalization, as the network is being trained to make consistent predictions from more points along the timeline.

Tab. 6(b) shows that it is essential to provide past predictions as the prefix. Without a prefix, the model is poor, underperforming our non-streaming baseline of only decoding at the end (30.630.6). This is because the model predicts duplicated predictions at each decoding point. Providing past captions outperforms this baseline substantially. We find that also adding previously predicted timestamps, does not improve over captions alone, suggesting that temporal boundaries of past events are not particularly informative for future events. Tab. 6(c) further shows that while training with a prefix, it is important to mimic inference behavior by including missing captions in earlier predictions.

Table 6(d) examines the choice of decoding points during inference. We find that using a stride, S ⁣= ⁣32S\!=\!32 (which equates to a decoding point at the middle and end of a 64-frame clip), performs considerably better than S ⁣= ⁣21S\!=\!21 (three uniformly chosen points). We observed qualitatively that using too many decoding points during inference can sometimes still result in the model making duplicate predictions (even with past prediction as the prefix). Whilst it is also possible to remove duplicate predictions with non-maximal suppression (NMS ), it is more challenging for captioning models as we also require a calibrated score for each event caption to perform NMS. We therefore leave an investigation of NMS for dense video captioning to future work.

2.3 Generalization to backbones and datasets

To show the generality of our method, we add our streaming modules onto both the GIT and Vid2Seq architectures, which we denote as Streaming GIT and Streaming Vid2Seq, respectively.

The last two rows of Tab. 4 shows that we improve substantially over both baselines consistently on three datasets. Our GIT baseline can process a maximum of Nf ⁣= ⁣16N_{f}\!=\!16 frames due to memory limitations, and our improvement also stems from the fact that we use Nf ⁣= ⁣64N_{f}\!=\!64 for our Streaming GIT thanks to our memory module. Vid2Seq pools visual tokens spatially such that it only uses a single token per frame. Therefore, our Streaming Vid2Seq does not use more frames than the baseline, and the improvement is due to our streaming of output event captions, which improves performance significantly as shown in Tab. 3.

Finally, we observe that our Vid2Seq baseline performs substantially better than the GIT baseline on YouCook2. This difference is due to the pretraining: We used the public, pretrained Vid2Seq checkpoint which was pretrained on YT-Temporal – a dataset with a similar domain to YouCook2. Note that this is the same experimental protocol as Vid2Seq , the current state-of-the-art.

3 State-of-the-art Comparison

Tab. 4 also compares our method to the state-of-the-art among dense video captioning methods using only video frames as inputs. We achieved substantial gains over prior, published works, notably improving CIDEr on ActivityNet by 11.0 points, and YouCook2 by 4.0 points, respectively. We also achieved improvements, albeit smaller, on SODA and Meteor. Our improvements on localization (F1) are smaller, showing the gains are more from better captioning qualities. Fig. 6 visualizes an example on ActivityNet.

We note that it is possible to further improve results, particularly on YouCook2, by using Automatic Speech Recognition (ASR) as an additional input modality . This is primarily because the spoken utterances of the actors are well aligned with the visual content, and because ASR was used in the annotation procedure of the dataset itself. However, integrating multiple modalities such as ASR is orthogonal to the main focus of this work.

Paragraph captioning. In addition to dense captioning, we also compare with state-of-the-art models on the same datasets for paragraph captioning, which aims to predict the captions throughout the entire video, but without any timestamps. Therefore, we only apply our streaming input model here, as the timestamps needed to assign decoding points during training are not available in this setting.

We train our Streaming GIT model for paragraph captioning, on both ActivityNet and Youcook2 . Tab. 5 shows that we achieve state-of-the-art results on this task too. The GIT baseline is our model trained on the same setting as our full model, but uses 16 input frames with all tokens concatenated for the decoder. This baseline already outperforms the state-of-the-art from visual-only inputs. Our model uses more input frames (64 frames), and further boosts the performance by 0.90.9 and 5.55.5 points on the two datasets, respectively, showing the benefits of our memory module which is consistent with our results in Tab. 1.

Conclusion and Future Work

We have proposed a streaming model for dense video captioning with two novel components: A clustering-based memory that can efficiently handle arbitrarily long videos with bounded computation, and a streaming decoding algorithm that enables our model to make predictions before the entire video has been processed. We achieve this streaming ability while also improving the state-of-the-art on five dense- and paragraph-captioning tasks.

Future work is to develop a benchmark for dense video captioning which requires reasoning over longer videos than current datasets, to better evaluate the abilities of streaming models such as ours.

Acknowledgments. We thank Chen Sun for helpful discussions.

References

Appendix

We provide further training details of our method in Sec. A.

Appendix A Training hyperparameters

All our experiments are conducted using the Scenic library and JAX . With the GIT architecture, we first pretrain on the WebLI dataset for general image captioning. WebLI contains 100M image-text pairs derived from alt-text from the internet. The image encoder is initialized from CLIP-L , and the language decoder is randomly initialized. During pretraining, we use the standard label-smoothed (factor 0.10.1) cross-entropy loss following GIT and train for 10 epochs. We use the Adam optimizer, with no weight decay. The learning rate is set to 5×10−55\times 10^{-5} with a batch-size of 10241024, with a cosine decay schedule. Following GIT , we use 0.2 ⁣×0.2\!\times lower learning rate for the image encoder.

When finetuning on dense-video captioning datasets , we freeze the image encoder. We again use the Adam optimizer with weight decay. We train for 20 epochs with batch size of 3232, and use a learning rate of 10−510^{-5}, dropped by 10 ⁣×10\!\times at the 16th epoch.

With Vid2Seq , we take the publicly released pretrained checkpointhttps://github.com/google-research/scenic/tree/main/scenic/projects/vid2seq, which is pretrained on the YT-Temporal dataset with a denoising and a captioning objective . When finetuning on dense-video captioning datasets , we follow their official training parameters. Specifically, we freeze the image encoder and pool the image tokens among the spatial dimensions to get one token per frame. The T5 decoder uses a dropout rate of 0.10.1. We again use Adam optimizer with weight decay. We train for 40 epochs with batch-size 32, and use a learning rate of 3×10−43\times 10^{-4} with a cosine decay schedule.

For all models, we follow the standard protocol to use beam-search decoding, with a beam size of 44 and a brevity penalty of 0.60.6 . We also emphasize that wherever applicable, all base architectures and backbones are consistent between comparisons and baselines.