UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Zhou Zhao, Xixin Wu, Helen Meng
Introduction
Audio generation is an important component of generative AI. Recently, the popularity of generative AI has induced increasingly emergent and varying needs in audio generation: audio is expected to be generated based on humans’s demands, such as speech synthesis (TTS), voice conversion (VC), singing voice synthesis (SVS), text-to-sound, and text-to-music. Prior works on audio generation tasks are commonly task-specific: their designs heavily leverage domain knowledge and their usage is restricted to fixed setups (Tan et al., 2021; Luo & Mesgarani, 2019; Zmolikova et al., 2023; Huang et al., 2021b; Cho et al., 2021). Instead of taking care of each task independently, this work is an attempt to achieve universal audio generation, which intends to accomplish multiple audio generation tasks with only one unified model. The universal audio generation model is expected to obtain sufficient prior knowledge in audio and related modalities, which has the potential to provide simple and effective solutions for the increasing needs of generating diverse types of audio.
The superiority of Large Languge Models (LLM) in text-generative tasks inspires a series of LLM-based models in audio generation (Wang et al., 2023a; Kharitonov et al., 2023; Huang et al., 2023b; Agostinelli et al., 2023; Borsos et al., 2023). Among these works, LLM’s capability in independent tasks has been extensively studied in tasks like text-to-speech (TTS) (Wang et al., 2023a; Kharitonov et al., 2023; Huang et al., 2023b) and music generation (Agostinelli et al., 2023; Copet et al., 2023), and achieves competitive performance. However, LLM’s ability to process multiple tasks with a unified model is less exploited in audio generation research: most existing LLM-based works are still designed for single tasks (Wang et al., 2023a; Kharitonov et al., 2023). We argue that achieving universality and versatility in audio generation through the LLM paradigm is promising but has not yet been comprehensively studied before this work.
Toward universal audio generation, this work presents UniAudio, which adopts LLM techniques and is able to generate multiple types of audio (speech, sounds, music, and singing) conditioned on various input modalities, such as phoneme sequences, textual descriptions, and audio itself. The proposed UniAudio is mainly featured as follows: First, all types of audio, along with all other input modalities, are tokenized as discrete sequences. Specifically, a universal neural codec model is built to effectively tokenize audio regardless of the audio type, and other tokenizers are used to tokenize other different modalites. Then, UniAudio concatenates the source-target pair as a single sequence. Lastly, UniAudio performs next-token prediction using LLM. The residual vector quantization (Zeghidour et al., 2021) based on neural codecs is used in the tokenization process, resulting in overly long token sequences (one frame corresponding to multiple tokens) that cannot be processed efficiently by LLM. A multi-scale Transformer architecture is designed to reduce computational complexity by modeling the inter- and intra-frame correlation separately. Specifically, a global Transformer module is used to model the inter-frame correlation (e.g. semantic level), and a local Transformer module is used to model the intra-frame correlation (e.g. acoustic level).
To demonstrate the scalability of UniAudio for new tasks, the building process of UniAudio takes two stages. Firstly, the proposed UniAudio is trained on multiple audio generation tasks jointly, which allows the model to obtain sufficient prior knowledge not only of the intrinsic properties of audio but also of the interrelationship between audio and other input modalities. Secondly, through fine-tuning, the trained model can seamlessly support more unseen audio generation tasks. Thus, UniAudio has the potential to become a foundation model for universal audio generation: it is able to continuously support emergent needs in audio generation. Experimentally, our UniAudio supports 11 audio generation tasks: the training stage includes 7 audio generation tasks, while 4 tasks are further added in the fine-tuning stage. The building process of UniAudio is scaled up to 165k hours of audio and 1B parameters. Among the 11 tasks, UniAudio consistently obtains competitive performance in both objective and subjective evaluations. State-of-the-art results are even achieved on most of these tasks. Further investigation suggests that training multiple tasks simultaneously in the training stage is mutually beneficial to each task involved. In addition, UniAudio can effectively adapt to new audio generation tasks and outperform task-specific models with a non-trivial gap.
To sum up, this work reveals that building universal audio generation models is necessary, promising, and beneficial. The main contributions of this work are summarized as follows: (1) Toward universal audio generation, UniAudio is presented as a unified solution for 11 audio generation tasks.
(2) Per methodology, UniAudio provides novel approaches for (i) sequential representations of audio and other input modalities; (ii) uniform formulation for LLM-based audio generation tasks; and (iii) efficient model architecture specifically designed for audio generation.
(3) Per experiments, the overall performance of UniAudio is well validated, and the benefits of building a versatile audio generation model are verified by exhaustive experimental results.
(4) Demo and code are released, in the hope that UniAudio can become a foundation model that supports emergent audio generation in future research.
UniAudio
This section introduces the technical details of the proposed UniAudio. Section 2.1 explains how audio and other modalities are tokenized. Then, all considered audio generation tasks are uniformly formulated in Section 2.2. Subsequently, the multi-scale Transformer architecture is proposed in Section 2.3 to handle the overly long sequence challenge caused by the adoption of neural codecs.
LLM are commonly used for sequential modeling, so audio and all other input modalities are tokenized before being processed. These processes for each modality are completed by independent modules. All of these modules are fixed in the optimization of UniAudio or parameter-free.
For all audio generation tasks considered in this work, audio, regardless of its types (speech, sounds, music, or singing), is the target to predict. Instead of modeling different types of audio separately, UniAudio intends to tokenize all types of audio as a single and unified modality (even though they commonly have distinct patterns, such as frequency span), which requires a model that is well-suited to mapping all audio types into a shared latent space. Following Wang et al. (2023a); Kharitonov et al. (2023), neural codec models (Défossez et al., 2022; Yang et al., 2023b; Kumar et al., 2023) are used in this work for audio tokenization. An audio signal of duration with sample rate can be represented by a sequence . An audio neural codec intends to compress and then recover it as using an encoder-decoder architecture with a quantization module:
where denotes the number of audio frames after down-sampling in the encoder, and denotes the feature dimension of the encoder. The discrete representations of audio are the intermediate product of the quantization process. Given any frame of hidden output , the integer vector is generated by Residual Vector Quantization (RVQ) (Zeghidour et al., 2021), where denotes the number of vector quantization layers. Iteratively, each element is the index among all pre-learned and fixed -th level quantizer vectors that has the smallest L2 distance to the residual between and the sum of all previous chosen quantizer vectors . With the discrete representation , is reconstructed as a close estimation of that can be used to recover with the decoder.
The discrete representation of all audio frames is a matrix and needs to be converted into a sequence before being processed by LM: it is simply flattened as a sequence, in which every element for one frame is consecutive. Without specifically stated, we set in our experiments. As the waveform can be recovered from with a neural codec decoder, the rest of this paper mainly discusses how to predict the audio token sequence using LLM techniques. As UniAudio intends to generate both speech and non-speech content, we build the codec model on our own and with broader data coverage. Details of our codec configuration is in Appendix E.
1.2 Other modalities
Besides audio, other modalities considered in UniAudio also need to be represented as sequences. In addition, most of these sequences are transformed into discrete ones through tokenization. The serialization and tokenization of these input modalities, along with their key features, are briefly summarized as below.
Phoneme: Phonemes are the basic units of speech pronunciation in linguistics. Phoneme sequences have multiple sources: (1) when only text is available, phoneme sequence without duration information can be obtained by text-to-phoneme mapping using a pronunciation dictionary; (2) when only speech is available, phoneme sequence with duration information is obtained by beam search of the DNN-HMM system (Hinton et al., 2012); (3) when both text and speech are available, phoneme sequence with duration information is obtained by forced alignment of the DNN-HMM system CMUDict (http://www.speech.cs.cmu.edu/cgi-bin/cmudict) is adopted as the pronunciation dict; kaldi recipe (https://github.com/kaldi-asr/kaldi/tree/master/egs/librispeech/s5/local/chain/run_tdnn.sh) is adopted to build the deep neural network-hidden Markov model (DNN-HMM) system..
MIDI: MIDI (Zhang et al., 2022) is widely used for singing voice synthesis tasks. F0 and duration information are included in the MIDI. We use the duration information to flatten the F0 sequence, so that the frame-level F0 sequence is obtained.
Text: Text acts as a effective carrier of human instructions in audio generation tasks (Yang et al., 2023a; Copet et al., 2023). In this work, these textual instructions are represented as continuous embeddings derived from pre-trained text LLM (Raffel et al., 2020), as these embeddings contain rich textual semantics. Processing these continuous embeddings with LLM is further clarified in Section 2.3The encoder of T5 (https://github.com/google-research/text-to-text-transfer-transformer) is used to extract the continuous text embeddings..
Semantic Token: The semantic tokens are derived from the continuous embeddings output by audio self-supervised learning (SSL) models. These continuous representations are highly informative and can be adopted in both speech understanding (Rubenstein et al., 2023) and generative tasks (Borsos et al., 2023). Following Huang et al. (2023b), these continuous representations are tokenized by performing K-means clustering (Hsu et al., 2021) over these continuous representations. Since the continuous representations are frame-level, the semantic tokens also encode duration informationThe 9-th layer hidden output of Hubert (Hsu et al., 2021) is adopted as the semantic token representations (https://github.com/facebookresearch/fairseq/hubert). The number of clusters for K-means is 500..
2 Unified Task Formulation
For all tasks considered in UniAudio, the target audio is generated based on given conditions. With the same target modality, i.e., audio, it is the conditions that define different audio generation tasks. However, even with the variance in conditions, all tasks can still be uniformly formulated as sequential modeling tasks that can be processed by LLM: both the target audio and the conditions are first transformed as sub-sequences and spliced as [conditions, target] sequences to be processed.
UniAudio supports 11 audio generation tasks in total. The sequential formats of each task are defined in Table 1, in which the sub-sequences of all modalities are derived as in Section 2.1. However, due to the unique configurations of each task, some of the condition sub-sequences are subject to task-specific pre-processing operations during the tokenization. For audio, these operations are mainly for data corruption, such as adding noise, reverberation, and speech mixed with other speakers in the raw audio before tokenization. For phoneme and semantic tokens, duration information is reserved by default but can also be removed. For singing voice synthesis and speech edit tasks, the duration information of phoneme is used. For TTS and I-TTS tasks, the duration information is not used. For MIDI, the duration information is used repeat the F0 sequence. For text embeddings, no operations are applied in this work.
To avoid ambiguity, some special discrete tokens (enclosed by <>) are inserted to indicate (1) the start and end of the whole sequence; (2) the start and end of each sub-sequence of a certain modality; and (3) the task identifier. For example, for a text-to-sound task sequence that generates target audio based on textual description, the whole sequence is like:
3 Multi-Scale Transformer
Previous work on LLM-based audio generation (Copet et al., 2023) advocates to modeling the discrete audio tokens as flattened sequences. If so, these sequences are processed in the length of , which is highly challenging considering the quadratic space complexity of Transformer (Vaswani et al., 2017) with respect to the lengths. Inspired by Yu et al. (2023), a multi-scale Transformer architecture is specifically designed for discrete audio sequences, which is a hierarchical model that processes the inter- and intra-frame correlation by global and local Transformer modules separately. An overview of the proposed architecture is in Figure 1. Instead of processing the whole flattened sequence token-by-token like prior works (Kharitonov et al., 2023), the multi-scale transformer considers patches (i.e., every consecutive token) as the global modeling units and then handles the tokens within each patch locally. Note that both the global and local Transformers are causal.
For audio token sequences, each patch accounts for consecutive audio tokens that exactly represent one audio frame. First, as suggested in Equation 2, regardless of the exact choices of each quantization vector , it is the summed quantization vector that is used to represent the audio frame. Thus, in the embedding stage, each patch (a.k.a., frame) is represented by the summed vector of the corresponding embeddings before entering the global Transformer. Second, the global Transformer is to predict audio frame-by-frame: to predict the frame , it outputs the continuous representations that include frame and all previous content. These continuous representations will be further processed by the local Transformer. Third, also as in Equation 2, given the hidden representation , the acquisition of is independent of any hidden output other than . Inspired by this, it is reasonable to predict the discrete tokens for frame , a.k.a., patch , only with the hidden output of global Transformer corresponding to frame . To be more detailed, as the acquisition of each token is auto-regressively dependent on its prior tokens , a local Transformer is adopted to predict the patch sequence in auto-regressive style. During this process, the corresponding vector output by the global transformer acts as a patch-level context, which is linearly transformed and then added to the embedded results of each token .
The proposed multi-scale Transformer architecture is also compatible with discrete and continuous sequences besides audio. For all discrete tokens except audio (phoneme, semantic, MIDI and special tokens), each token has independent semantics and thus should account for one patch. So these discrete tokens repeat for times to fill each patch. The continuous text embeddings are also repeated for times for the same purpose. Additionally, their embedding process is replaced by a linear transformation while their predicting targets for local Transformer are consecutive special tokens
The design of the proposed multi-scale Transformer can effectively reduce computational complexity. First, the equivalent sequence length for the global Transformer is reduced from to , which makes the global modeling cost independent to and thus the adoption of a larger becomes feasible. Second, the intra-patch computation to generate the discrete tokens for each frame is offloaded to the local Transformer. The computation on the local transformer is comparatively light since it only processes the very short sequence (fixed to the length of ) and empirically has fewer parameters than the global Transformer by design.
Experiments
This section first introduces the experimental setup in Section 3.1. The results for the training stage and the fine-tuning stage are presented in Section 3.2 and 3.3 respectively. Ablation studies are presented in Section 3.4.
Data and Model: UniAudio is built on labeled datasets. Specifically, 12 datasets are adopted in this work, all of which are publicly available. The overall audio volume is 165K hours. Detailed data statistics and their adoption for each task are in Appendix A.1. Discrete tokens from all modalities form a joint vocabulary of size 4212, including all special tokens. Vanilla Transformer decoder layers with causality are consistently adopted in global and local Transformer. The overall parameter budget is roughly 1B. Detailed model configuration is in Appendix A.2. Existing neural codec models are sub-optimal for universal audio generation, mainly due to data coverage. An improved neural codec model is then built with fewer quantization levels , smaller frame-per-second rate, higher quality, and wider coverage (see Appendix E).
Training and Inference: The training stage includes 7 tasks while 4 new tasks are added in the fine-tuning stage. Table 1 specifies the tasks for fine-tuning only. Both the training and fine-tuning are completed with 16 AMD MI200-64G GPUs. The detailed configuration of optimization is in Appendix A.3. To retain the performance of previous tasks during fine-tuning, following Conneau et al. (2020), the training data are re-sampled with respect to tasks with . Top-k sampling is adopted consistently for inference, in which and the temperature are set to 30 and 0.8, respectively. As the global Transformer does not directly predict tokens, the sampling process only happens in the local Transformer inference.
Evaluation: For evaluation, most tasks are evaluated using both objective and subjective metrics Following the setting of DiffSinger (Liu et al., 2022), SVS tasks don’t report the objective results. Generally, for objective evaluation, Word Error Rate (WER) is used to evaluate the intelligibility of generated speech; Similarity Score (SIM) is for similarity in terms of speaker identityWER and SIM evaluation models follow Wang et al. (2023a); Perceptual Evaluation of Speech Quality (PESQ), VISQOLhttps://github.com/google/visqol, DNSMOS https://github.com/microsoft/DNS-Challenge/tree/master/DNSMOS and Mel Cepstral Distortion (MCD) are signal-level quality metrics derived from human auditory research; Following (Copet et al., 2023), Fréchet Audio Distance (FAD), Kullback-Leiber (KL) Divergence, and Fréchet Distance (FD) are for audio fidelity and audio similarity; For subjective evaluation, MOS and SMOS are adopted to provide human-centric judgment for speech and sing related tasks. For text-to-sound and text-to-music tasks, we use overall quality (OVL), and relevance to the text input (REL) (Copet et al., 2023). Note all subjective results are obtained from Amazon Mechanical Turkhttps://www.mturk.com/ for fair comparison. Appendix F shows details of the subjective evaluation process.
2 The Results of 7 generative tasks in the training stage
This section presents the overall evaluation results of the proposed UniAudio model over all 7 audio generation tasks during the training stage. A comprehensive comparison is conducted between UniAuduio and multiple prior works on each task, including not only the LM-based methods but also the diffusion model-based methods as well as other conventional audio generation methods. The detailed comparison is presented in Appendix B. We selected one of the most advanced prior work in each task and present the results in Table 2.
As suggested in Table 2, UniAudio is a versatile system that can handle all 7 audio generation tasks together and achieve competitive performance. Per subjective evaluation, UniAudio surpasses the baselines in 3 out of 6 tasks (TTS, VC, Sound); per objective evaluation, it achieves better results on 5 out of the 7 tasks except SVS and Music. We also find UniAudio under-perform on several metrics. UniAudio’s subjective performance for SE and TSE is less competitive compared with its competitors, which is also observed in previous literature (Erdogan et al., 2023) that the signal-level evaluation metrics may not be suitable for LM-based generative methods. UniAudio cannot surpass the selected competitor (Copet et al., 2023) in the Text-to-Music task. We note that (Copet et al., 2023) is built with more private labeled data than our UniAudio.
3 The Results of 4 generative tasks in the fine-tuning stage
As UniAudio is designed to continuously support new audio generation tasks, this section reports UniAudio’s performance on unseen tasks. The model is obtained by fine-tuning over 4 new tasks jointly and the results are presented in Table 3. Similar to section 3.2, for each task, we compare UniAudio’s performance with one selected prior work and report the detailed results in Appendix B.
As shown in Table 3, the fine-tuned UniAudio model surpasses its baselines in audio edit and speech dereverberation and is approaching the ground-truth quality in the Instructed TTS task. For speech editing, UniAudio shows considerable improvement compared to generating the whole sentence.
4 Ablation Study
To further validate our claim that building a unified model for all 11 audio generation tasks is promising and beneficial, more ablation studies are conducted. In Appendix C.1, we demonstrate that the joint-trained UniAudio model consistently outperforms the models that are trained for each specific taskNote the task-specific models are built with the corresponding subset of the training data., regardless they are included in the training stage or the fine-tuning stage. In Appendix C.2, we additionally validate that fine-tuning over the 4 new audio generation tasks does not affect UniAudio’s performance on the original 7 tasks. In Appendix C.3, we observe that UniAudio can consistently benefit from increased training data volume of each task, which provides another reason to build universal audio generation models: these models are easier to scale up as the data collection is more feasible. We provide more discussion in Appendix D about the effectiveness of building a universal audio generation model.
4.2 The effectiveness of multi-scale transformer model
As in section 2.3, the adoption of neural codecs has become a popular choice of LLM-based audio generation but causes an overly long sequence issue that needs further consideration. This section compares the proposed multi-scale Transformer with four representative approaches in this field: Flattening Prediction (e.g. SPEARTTS (Kharitonov et al., 2023)), Coarse first prediction (e.g. VALL-E (Wang et al., 2023a)), Parallel prediction (e.g. AudioGen (Kreuk et al., 2022)), and Delay prediction (e.g. MusicGen (Copet et al., 2023)). Figure 2 illustrates the prediction order of these five architectures. Experiments are conducted on text-to-speech and text-to-music tasks and the results are reported in Table 4 and 5 respectively Results are based on unofficial implementations..
Auto-Regression and Performance: Among all 4 baselines aforementioned, Copet et al. (2023) claims that the flattening method provides the best audio generation quality. they further claim that the superior performance of flattening prediction is mainly attributed to the auto-regressive property; the other three methods do not reserve this property as the concurrent prediction is introduced (see Fig. 2). Under the scenario of codec adoption, we reinterpret the auto-regressive property as: current token prediction is based on all tokens of previous frames and the previous tokens within the current frame, or formally, the prediction of the current token is based on tokens: . With this definition, we claim that the proposed multi-scale transformer is also auto-regressive.
Aligned with Copet et al. (2023), our experiments also validate the importance of the auto-regressive property. As in Table 4 and 5, flattening prediction brings better generation quality than parallel, coarse first, and delay prediction. Additionally, with the same auto-regressive property, our proposed multi-scale transformer achieves a comparable performance with flattening prediction in terms of generation quality, which, again, validates the importance of auto-regression.
Efficiency: Besides generation quality, efficiency is a major concern of audio generation. Although with the auto-regressive property, the flattening prediction is sub-optimal in terms of efficiency: the modeling is based on the long sequence, which has a space complexity of in self-attention. As increasing gives higher reconstruction quality at the cost of longer sequences and more computation, this issue becomes more severe when a larger is adopted. Since the sequence length grows proportionally with , we experimentally find it difficult to train with . By contrast, the proposed multi-scale Transformer distributes the inter- and intra-frame modeling to the global and local sub-modules respectively, which thus alleviates the space complexity to . Finally, without the requirement of auto-regression, methods like parallel, coarse first, and delay predictions achieve better efficiency due to the adoption of concurrent predictions. Since the space complexity is independent to , training a larger with the multi-scale transformer is then feasible.
Experimentally, the proposed multi-scale transformer considerably reduces the time and memory cost compared with the flatting prediction. It still costs more time and memory compared with the other three baselines.
Based on the observations above, we claim that the proposed multi-scale transformer is an auto-regressive architecture that achieves a better trade-off between generation quality and efficiency.
Related Works
This work is an attempt to achieve universal audio generation through LLM-based techniques. There is a long research history for many audio generation tasks. Conventionally, the design of these tasks heavily leverages the domain knowledge of each specific task, and their workflows are distinctive from each other: For tasks like TTS, SE, TSE, TT-Music, VC, S-Edit, SD, SVS, (1) their neural network architectures are based on Transformer (Ren et al., 2020) or others (Oord et al., 2016; Luo & Mesgarani, 2019); (2) their training objectives can be either in time-domain (Luo & Mesgarani, 2019), frequency-domain (Yu et al., 2017) or others (Gu et al., 2021; Shen et al., 2023); (3) their designs are inspired by and derived from linguistics and phonetics (Zen et al., 2013), signal processing (Griffin & Lim, 1984), auditory perception (Shadle & Damper, 2001) and machine learning (Wang et al., 2016) research, etc; (4) they use different generative models, such as diffusion model (Shen et al., 2023; Wang et al., 2023b), flow (Le et al., 2023), Seq2Seq (Ren et al., 2020; Liu et al., 2021).
The prosperity of LLM techniques (Radford et al., 2019; OpenAI, 2023) significantly promotes progress in audio generation research in several directions. First, the large language models, along with the prompt methods, inspired multiple emergent audio generation tasks that are based on textual instruction or descriptions from humans, such as Instruct-TTS (Yang et al., 2023a), Text-to-sound (Kreuk et al., 2022; Huang et al., 2023a) and text-to-music Copet et al. (2023); Agostinelli et al. (2023). Second, besides the text, audio can also be tokenized as discrete sequences (Zeghidour et al., 2021; Défossez et al., 2022; Kumar et al., 2023) that can be further processed by LMs. LM-based audio generative models then show superior capability in generalization towards unseen speakers (Wang et al., 2023a), low resources (Kharitonov et al., 2023) and multilingual (Zhang et al., 2023) scenarios. These methods also achieve state-of-the-art results in overall performance within their own scopes. Finally, the LM-like model can be further combined with existing generative models (e.g., diffusion models Rombach et al. (2022)) to obtain improved generation quality.
It is laborious to handle each audio generation task case-by-case, especially when considering the data shortage as well as the emergent and varying needs in this area. Alternatively, building a universal audio generation model is a promising and practical paradigm. Given the rapid progress in audio generation research, recent designs of audio generation, including LM-based ones, tend to support multiple audio generation tasks simultaneously. Some pioneer works (Wang et al., 2023c; Le et al., 2023; Shen et al., 2023; Liu et al., 2023b; Jiang et al., 2023) clearly consider supporting multiple tasks as a key strength; the designs of other prior works (Borsos et al., 2023; Kharitonov et al., 2023; Shen et al., 2023) do show the potential to generate audio in a broader sense than what they originally claim. Following these pioneering research works, UniAudio supports an extended coverage of 11 audio generation tasks in a unified LM-based model.
Limitation
Not all known audio generation tasks are included in the proposed UniAudio, such as noise removal, noise speech edit (Wang et al., 2023c) and speech-to-speech translation (Rubenstein et al., 2023; Barrault et al., 2023). All new tasks added in fine-tuning are formulated with the known modalities in the training stage; Introducing new modalities during fine-tuning is unexplored in this work. Current UniAudio considers neither unlabeled data nor domain-specific foundation models, which can possibly further improve the overall performance. The samples generated by UniAudio are not guaranteed in quality and may contain errors.
Conclusion
To handle the emergent and varying needs in audio generation, this work is an attempt to achieve universal audio generation. UniAudio is proposed as a unified LM-based generative model that supports 11 different audio generation tasks. In experiments, the proposed UniAudio provides competitive performance on all 11 tasks. It also empirically demonstrates the capability of continuously integrating unseen audio generation tasks. Demo and code are released, in the hope that UniAudio can become a foundation model for universal audio generation in further research.
Ethical Statement
We are delving into the revolutionary field of generating diverse audio using large language model techniques. We find ourselves at the confluence of innovation and responsibility. It is imperative to acknowledge the ethical dimensions of our work and ensure that our contributions are employed for the betterment of society.
Being Open: As we advance in this domain, it’s crucial to ensure that the benefits of this technology are widespread and not limited to a privileged few. Our code is released publicly along with this submission to ensure equal access for each person. All experiments are based on open-accessible datasets that allow research-oriented comparison and reproduction.
Avoid Misuse: While our model can produce a myriad of audio content ranging from music to speech, there’s potential for misuse in the generation of misinformation, deepfake audio, or any harmful content. We advocate for adopting our code and model responsibly, with full respect to individual privacy and observance of regulations. Concerning the potential misuse of our model, checkpoints will not be released.
References
Appendix A Experimental Setup
This appendix describes experimental setups in detail, including data statistics, model architecture and optimization strategy.
12 public datasets are adopted in this work for training. Besides, several test sets are additionally used only for zero-shot evaluation. The statistics of these datasets are in Table 6. Datasets adoption for each task is described in Table 7. Note some datasets are adopted by more than one task.
A.2 Model Configuration
The model configuration of the proposed multi-scale Transformer is described in Table 8.
A.3 Optimization
The optimization configurations adopted in both the training and fine-tuning stages are presented in Table 9
Appendix B The Details of Experiments
This section presents detailed experimental results on each task. In the following, if the training set and test sets come from different datasets, we label them as zero-shot settings.
For TTS tasks, UniAudio is compared with the many previous SOTA models, Table 10 presents the results. For FastSpeech 2, we only conduct QMOS evaluation as its implementation adopts speaker id as input https://github.com/ming024/FastSpeech2. We can see that UniAudio obtains better performance in terms of WER, SIM than YourTTS, VALL-E, NaturalSpeech 2 and Make-A-Voice. Compared with VoiceBox, UniAudio also gets comparable performance in terms of objective metrics. From the MOS evaluation, we can see that UniAudio can generate high-quality speech compared with previous SOTA works. Furthermore, UniAudio realizes the best zero-shot clone ability (e.g. SMOS is 3.56 and SIM is 0.708). More experiments, such as cross-lingual zero-shot TTS and Mandarin Chinese speech synthesis can be found in demo page. For VC task, we conducted experiments on VCTK dataset, we randomly chose 200 audio pairs. PPG-VC and YourTTS are trained on small-scale datasets. Make-A-Voice and LM-VC We seek help from the authors, they provide the inference results. are trained on large-scale datasets as the same as UniAudio. Compared with previous work, UniAudio got better performance in voice conversion tasks.
B.2 Speech Enhancement and Target Speaker Extraction
For the SE task, we compare with previous SOTA methods, including discriminative methods (such as FullSubNet and FullSubNet+) and generative methods (such as SGMSE+ and NADiffuSE). Note that the CDiffuSE and NADiffuSE are both trained on the voicebank-demand dataset. Other models never saw the VCTK dataset in the training stage. We obtain the inference results based on their open-source models. Table 11 presents the results, we can see that UniAuido obtains the best DNSMOS score. The PESQ and VISQOL scores are lower than other SOTA methods, we think these metrics may not accurately assess the performance of generative methods. The similar finding is also observed in previous literature (Erdogan et al., 2023) that the signal-level evaluation metrics may not be suitable for generative methods. In contrast, we recommend using DNSMOS and MOS scores as the main metrics. UniAuido can get good results in extremely noisy environments, we recommend readers refer to the demo page. For the TSE task, we conducted experiments on the LibriMix test set. The popular TSE systems: VoiceFilter https://github.com/Edresson/VoiceSplit and SpeakBeamhttps://github.com/BUTSpeechFIT/speakerbeam are used as baseline systems. As Table 11 shows, we can see that UniAudio obtains the best performance in terms of DNSMOS and MOS.
B.3 Singing Voice Synthesis
Following Make-A-Voice, we conduct experiments on the M4Singer test set. We compare the generated singing samples with other systems, including 1) Diffsinger; 2) Make-A-Voice, a two-stage audio language model for singing voice generation. As illustrated in Table 12, we can see that UniAudio gets comparable results with Make-A-Voice and Diffsinger.
B.4 Text-to-sound and text-to-music generation
The text-to-sound generation task has attracted great interest in audio research. Following Diffsound (Yang et al., 2023c), most of the methods evaluate their systems on the AudioCaps (Kim et al., 2019) test set. However, we found that if the training data includes the AudioCaps data, the model is easy to overfit with AudioCaps. As a result, the best performance can be obtained when the model only trains on the Audiocaps. In this study, we conduct a zero-shot evaluation on the Cloth test set (Drossos et al., 2020). Table 13 shows the results. We can see that UniAudio obtains better performance than Diffsound and AudioLDM. Compared to recent SOTA models, such as Tango and Make-an-Audio 2, UniAudio also gets comparable performance. For the text-to-music task, we follow MusicGen (Copet et al., 2023), evaluating our methods on MusicCaps (Agostinelli et al., 2023). Compared with previous SOTAs, UniAudio gets a comparable performance with other models. From the MOS evaluation performance, we can see that MusicGen is better than our current models. We speculate one of the reasons is that MusicGen uses a large-scale high-quality dataset (20k hours).
B.5 Audio Edit
Audio edit aims to edit the original audio based on Human’s instruction. AUDIT (Wang et al., 2023d) is the SOTA model in audio edit task, which designs a data simulation strategy to get triplet training and test data (e.g., {audio, audio, text}). The authors set 5 different tasks, including adding, dropping, replacing, inpainting and super-resolution, and simulated large-scale data for each task. To validate that our pre-trained model can be fine-tuned with small-scale data, we choose adding, dropping and super-resolution tasks to fine-tune simultaneously. To finish the fine-tuning process, we define a new task label: Audit_task. The experimental results as Table 14 shows. We can observe that: (1) UniAudio can get better performance with the previous SOTA model. (2) Fine-tuning pre-trained UniAudio can get better performance than training it from scratch, which further validates the effectiveness of pre-training a model on large-scale training data.
B.6 Instructed TTS
Using instruction to guide speech synthesis has received great attention (Guo et al., 2023; Yang et al., 2023a). In this part, we fine-tune the UniAudio model on the PromptSpeech (Guo et al., 2023) dataset. Furthermore, we also try to train a UniAudio model from scratch with the PromptSpeech dataset. Different from previous works that designed special style encoders to capture the style information from text descriptions, we directly use the T5 text encoder to extract representations from text and then combine it with the phoneme sequence input to the UniAudio, which is more convenient.Note that the authors of PromptTTS (Guo et al., 2023) told us their objective metrics tools, checkpoints, and generated samples have been lost due to the machine errors. Thus we cannot fairly compare with them. Table 15 shows the results, we can see that UniAudio has good performance in terms of style control and speech quality when compared with the ground truth samples.
B.7 Speech Dereverberation
For the speech dereverberation task, we use the Room Impulse Response (RIR) data from the openSLR26 and openSLR28 dataset, and the speech data from the LibriTTS clean part. We simulate about 100 hours of training data and 1 hour of test data. We compare with previous SOTA systems, such as FullSubNet, FullSubNet+ and SGMSE+. Table 16 presents the results. We can see that UniAudio obtains the SOTA performance in speech dereverberation tasks with small-scale training data in terms of DNSMOS metric. Similar with speech enhancement task, we speculate that PESQ may not suitable for the generative methods.
B.8 Speech Edit
For the speech edit task, we use the LibriTTS dataset. In practice, we randomly choose some words to mask in the training stage. We expect the model to recover the whole speech based on the phoneme sequence. In the inference stage, we can mask the region that we want to update in the speech and input the new words so that the model can edit the speech. For this task, we take the TTS system that regenerates a complete waveform from the whole sentence to be edited as the baseline. In the evaluation, we mainly validate three situations: (1) word replacement; (2) insert a new word; and (3) delete a word. For each situation, we randomly chose 10 sentences from the LibriTTS test clean set.
Appendix C Ablation study
In this part, we explore whether multi-task training can bring better performance than task-specific training. To answer this question, we use the same model trained on different tasks, respectively. Table 17 shows the experimental results, UniAudio (single) means that the model is trained on a single task. We observe that multi-task training brings the gain over all of the tasks. In Appendix D, we give some potential reasons why multi-task training can bring improvement.
C.2 Fine-tuning the pre-trained model on the new task will influence the performance on previous tasks?
In this part, we conduct experiments to explore whether fine-tuning the pre-trained model on new tasks will influence the performance of previous tasks. We evaluate the pre-trained UniAudio model (trained on 7 tasks) and fine-tuned UniAudio model (fine-tuned on 4 new tasks) on 7 tasks. Figure 3 shows the results. We can see that the performance does not significantly drop on previous training tasks, which demonstrates that UniAudio has the potential to add new tasks continuously without losing previous task knowledge.
C.3 The influence of data quantity
In this part, we conduct experiments to explore the influence of data quantity, we give three settings: (1) using all of the data; (2) using training data for each task; (3) using training data for each task. We present the results in Figure 4. Based on the experimental results, this work claims that the data quantity is a key point to building a strong audio foundation model. In the future, we will explore to use of more unlabeled data to help improve the performance.
Appendix D Why UniAudio Can Work Well?
From the previous discussions, we can see that the universal modeling strategy brings improvement for different tasks. In this part, we try to give some potential explanations.
(1) Deterministic latent space: we formulate different modalities into a deterministic latent space (fixed vocabulary) by tokenization. Different tokens can be seen as specific ’words’, and we can use a next-token prediction strategy to train the model. Similar to GPT-series (Radford et al., 2018; 2019), such strategy creates the opportunity for the model to learn the intrinsic properties of audio and the interrelationship between audio and other modalities.
(2) Shared information between different types of audio: Although multiple types of audio (speech, sounds, music, and singing) present significant differences in the time domain or frequency domain, neural audio codec models effectively capture their shared information (rethinking the working principle of neural codecs, which similar information will be allocated the same token id). Due to the shared information that exists in different types of audio, multi-task training can be seen as increasing training data for each task.
(3) Data augmentation perspective: We speculate that multi-task training can be viewed as data augmentation for some tasks. Considering the TTS and VC task’s definition: TTS:
Appendix E The details of Audio Codec Models
In this part, we give more details about our neural audio codec model in Section 2.1.1. We adopt a similar encoder-decoder framework with the Encodec model, the difference includes: (1) we replace the multi-scale STFT-based (MS-STFT) discriminator as our multi-scale Mel-based discriminator. (2) We rewrite the vector quantization implementation Please refer to our source code to find the details. based on Encodec’s open-source version https://github.com/facebookresearch/encodec/blob/main/encodec/quantization/core_vq.py, making it more suitable for DDP training. Figure 5 shows the details of the mel-based discriminator. We combine the mel-spectrogram and log-mel-spectrogram features and then input them into a network consisting of several convolutional layers. Our motivation is that the mel-spectrogram has a strong intrinsic inductive bias, especially for sounds and music-related audio (the SOTA sounds or music classification systems are based on the log-mel-spectrogram in the literature.). Thus, we speculate that choosing a mel-spectrogram-based discriminator can better promote high-fidelity audio reconstruction. In our experiments, we use 6 different discriminators with different configurations In our experiments, we find the mel-based discriminator brings better reconstruction performance when we train a universal neural audio codec.. Specifically, we set the hidden_dim as {64, 128, 256, 512, 512, 512} and the hop length as {32, 64, 128, 256, 512, 1024}. We train the neural audio codec model based on the Librilight and AudioSet datasets. Table 18 demonstrates that the neural codec model adopted in this work outperforms prior Encodec (Défossez et al., 2022).
Appendix F Subjective Evaluation
For TTS and VC tasks, we focus on speech quality (QMOS) and speaker similarity (SMOS). The details are as follows. For speech quality evaluation, we conduct the MOS (mean opinion score) tests and explicitly ask the raters to focus on examining the audio quality and naturalness, and ignore the differences of style (timbre, emotion, and prosody. The testers present and rate the samples, and each tester is asked to evaluate the subjective naturalness on a 1-5 Likert scale.
For speaker similarity evaluation, we ask the raters to focus on the similarity of the speaker identity (timbre) to the reference, and ignore the differences in content, grammar, or audio quality. We paired each synthesized utterance with a reference utterance to evaluate how well the synthesized speech matched that of the target speaker.
For SE and TSE tasks, we write explicit instructions to ask the rater to assess the generated speech. Refer to Figure 6 to see the details.
For SVS, we also conduct quality MOS (QMOS) and style similarity MOS (SMOS). Different from TTS’s SMOS evaluation, we explicitly instruct the raters to focus on the similarity of the style (timbre, emotion, and prosody) to the reference, and ignore the differences in content, grammar, or audio quality.
For sound and music generation tasks, we follow AudioGen (Kreuk et al., 2022) and MusicGen (Copet et al., 2023) to evaluate (1) overall quality (OVL), and (2) relevance to the text input (REL).
Our subjective evaluation tests are crowd-sourced and conducted by 20 native speakers via Amazon Mechanical Turk. The screenshots of instructions for testers have been shown in Figure 6. We paid about $500 on participant compensation. A small subset of speech samples used in the test is available at https://uniaudio666.github.io/demo_UniAudio/.