A Complete Survey on Generative AI (AIGC): Is ChatGPT from GPT-4 to GPT-5 All You Need?

Chaoning Zhang, Chenshuang Zhang, Sheng Zheng, Yu Qiao, Chenghao Li, Mengchun Zhang, Sumit Kumar Dam, Chu Myaet Thwal, Ye Lin Tun, Le Luang Huy, Donguk kim, Sung-Ho Bae, Lik-Hang Lee, Yang Yang, Heng Tao Shen, In So Kweon, Choong Seon Hong

Introduction

Generative AI (AIGC, a.k.a AI-generated content) has made headlines with intriguing tools like ChatGPT or DALL-E (Ramesh et al., 2021), suggesting a new era of AI is coming. Under such overwhelming media coverage, the general public are offered many opportunities to have a glimpse of AIGC. However, the content in the media report tends to be biased or sometimes misleading. Moreover, impressed by the powerful capability of ChatGPT, many people are wondering about its limits. Very recently, OpenAI released GPT-4 (OpenAI, 2023) which demonstrates remarkable performance improvement over the previous variant GPT-3 as well multimodal generation capability like understanding images. Impressed by the powerful capability of GPT-4 powered by AIGC, many are wondering about its limits: can GPT-5 (or other GPT variants) help next-generation ChatGPT unify all AIGC tasks? Therefore, a comprehensive review of generative AI serves as a groundwork to respond to the inevitable trend of AI-powered content creation. More importantly, our work comes to fill this gap in a timely manner.

The goal of conventional AI is mainly to perform classification (Lu and Weng, 2007) or regression (Lewis-Beck and Lewis-Beck, 2015). Such a discriminative approach renders its role mainly for analyzing existing data. Therefore conventional AI is also often termed analytical AI. By contrast, generative AI differentiates by creating new content. However, generative AI often also requires the model to first understand some existing data (like text instruction) before generating new content (Brown et al., 2020; Ramesh et al., 2022). From this perspective, analytical AI can be seen as the foundation of modern generative AI and the boundary between them is often ambiguous. Note that analytical AI tasks also generate content. For example, the label content is generated in image classification (Krizhevsky et al., 2012). Nonetheless, image recognition is often not considered in the category of generative AI because the label content has low dimensionality. Typical tasks for generative AI involve generating high-dimensional data, like text or images. Such generated content can also be used as synthetic data for alleviating the need for more data in deep learning (He et al., 2022b). An overview of the popularity of generative AI as well as its underlying reasons, is presented in Sec.2.

As stated above, what distinguishes generative AI from conventional one lies in its generated content. With this said, generative AI is conceptually similar to AIGC (a.k.a. AI-generated content) (of Information and Technology, 2022). In the context of describing AI-based content generation, these two terms are often interchangeable. In this work, we call the content generation tasks AIGC for simplicity. For example, ChatGPT is a tool for the AIGC task termed ChatBot (Caldarini et al., 2022), which is the tip of the iceberg considering the variety of AIGC tasks. Despite the high resemblance between generative AI and AIGC, these two terms have a nuanced difference. AIGC focuses on the tasks for content generation, while generative AI additionally considers the fundamental technical foundations that support the development of various AIGC tasks. In this work, we divide those underlying techniques into two classes. The first class refers to the generative modeling techniques, like GAN (Goodfellow et al., 2014) and diffusion model (Ho et al., 2020), which are directly related to generative AI for content creation. The second class of AI techniques mainly consists of backbone architecture (like Transformer (Vaswani et al., 2017)) and self-supervised pretraining (like BERT (Devlin et al., 2018) or MAE (He et al., 2022a)). Some of them are developed in the context of analytical AI. However, they have also become essential for demonstrating competitive performance, especially in challenging AIGC tasks. Considering this, both classes of underlying techniques are summarized in Sec.3.

On top of these basic techniques, numerous AIGC tasks have become possible and can be straightforwardly categorized based on the generated content type. The development of various AIGC tasks is summarized in Sec.4, Sec.5 and Sec.6. Specifically, Sec.4 and Sec.5 focus on text output and image output, respectively. For text generation, ChatBot (Caldarini et al., 2022) and machine translation (Yang et al., 2020a) are two dominant tasks. Some text generation tasks also take other modalities as the input, for which we mainly focus on image and speech. For image generation, two dominant tasks are image restoration and editing (Liu et al., 2022c). More recently, text-to-image has attracted significant attention. Beyond the above two dominant output types (i.e. text and image), Sec.6 covers other types of output, such as Video, 3D, Speech, etc.

As technology advances, the AIGC performance gets satisfactory for more and more tasks. For example, ChatBot used to be limited to answering simple questions. However, the recent ChatGPT has been shown to understand jokes and generate code under simple instruction. Text-to-image used to be considered a challenging task; however, recent DALL-E 2 (Ramesh et al., 2022) and stable diffusion (Rombach et al., 2022) have been able to generate photorealistic images. Therefore, opportunities of applying the AIGC to the industry emerge. Sec.7 covers the application of AIGC in various industries, including entertainment, digital art, media/advertising, education, etc. Along with the application of AIGC in the real world, numerous challenges like ethical concerns have also emerged and they are disused in Sec.8. Alongside the current challenges, an outlook on how generative AI might evolve is also presented.

Overall, this work conducts a survey on generative AI through the lens of generated content (i.e. AIGC tasks), covering its underlying basic techniques, task-wise technological development, application in the industry as well as its social impact. An overview of the paper structure is presented in Figure 4.

Overview

Adopting AI for content creation has a long history. IBM made the first public demonstration of a machine translation system at its head office in New York in 1954. The first computer-generated music came out with the name “Illiac Suite” in 1957. Such early attempts and proof-of-concept successes caused a high expectation of the AI future, which motivated governments and companies to invest numerous resources in AI. Such a high boom in investment, however, did not yield the expected output. After that, a period called AI winter came, which dramatically undermines the development of AI and its applications. Entering the 2010s, AI has again become popular again, especially after the success of AlexNet (Krizhevsky et al., 2012) for ImageNet classification in 2012. Entering the 2020s, AI has entered a new era of not only understanding existing data but also creating new content (Brown et al., 2020; Ramesh et al., 2022). This section provides an overview of generative AI by focusing on its popularity and why it gets popular.

A good indicator of ‘how popular a certain term is’ refers to search interest. Google provides a promising tool to visualize search frequency, called Google trends. Although alternative search engines might provide similar functions, we adopt Google trends because Google is one of the most widely used search engines in the world.

Interest over time and by region. Figure 1 (left) shows the search interest of generative AI, which indicates that the search interest significantly increased in the past year, especially after October 2022. Entering 2013, this search interest reaches a new height. A similar trend is observed for the term AIGC, see Figure 2 (left). Except for interest over time, Google trends also provides region-wise search interest. The search heatmaps for generative AI and AIGC are shown in Figure 1 (right) and Figure 2 (right), respectively. For both terms, the main hot regions include Asia, Northern America, and Western Europe. Most notably, for both terms, China ranks highest among all countries with a search interest of 100, followed by around 30 in Northern America and 20 in Western Europe. It is worth mentioning that some small but tech-oriented countries also have a very high search interest in generative AI. For example, the three countries that rank top on the country-wise search interest are Singapore (59), Israel (58), and South Korea (43).

Generative AI v.s. AIGC. Figure 3 shows a comparison between generative AI and AIGC for the search interest. Here, we define the interest ratio of generative AI and AIGC as GAI/AIGC. A major observation is that China prefers to use the term AIGC compared with generative AI with the GAI/AIGC ratio being 15/85. By contrast, the GAI/AIGC in the US is 90/10. In many countries, including Russia and Brazil, the GAI/AIGC is 100/0. Overall, most countries prefer generative AI to AIGC, which makes generative AI have an overall higher search interest than AIGC. The reason that China becomes the leading country to adopt the term AIGC is not fully clear. A possible explanation is that AIGC is shortened to a single word and thus is easier to use. We also search the Chinese version of generative AI and AIGC on Google trends, however, the current demonstration is not sufficient.

2. Why does it get popular?

The recent surging interest in generative AI in the last year can be mainly attributed to the emergence of intriguing tools like Stable diffusion or ChatGPT. Here, we discuss why generative AI gets popular by focusing on what factors contributed to the advent of such powerful AIGC tools. The reasons are summarized from two perspectives: content need and technology conditions.

The way we communicate and interact with the world has been fundamentally changed by the Internet, for which digital content plays a key role. Over the last few decades, the content on the web has also undergone multiple major changes. In the Web 1.0 era (the 1990s-2004), the Internet was primarily used to access and share information, with websites mainly static. There was little interaction between users and the primary mode of communication was one-way, with users accessing information but not contributing or sharing their own content. The content was largely text-based and it was mainly generated by professionals in the relative fields, like journalists generating news articles. Therefore, such content is often called Professional Generated Content (PGC), which has been dominated by another type of content, termed User Generated Content (UGC) (Paulussen and Ugille, 2008; Koolen et al., 2013; Timoshenko and Hauser, 2019). In contrast to PGC, UGC in Web 2.0 (O’reilly, 2009) is mainly generated by users on social media, like Facebook (Kim and Johnson, 2016), Twitter (Liu et al., 2017), Youtube (Holland, 2016), etc. Compared with PGC, the volume of UGC is significantly larger, however, its quality might be inferior.

We are currently transitioning from Web 2.0 to Web 3.0 (Rudman and Bruwer, 2016). With defining features of being decentralized and intermediary-free, Web 3.0 also relies on a new content generation type beyond PGC and UGC to address the trade-off between volume and quality. AI is widely recognized as a promising tool for addressing this trade-off. For example, in the past, only those users that have a long period of practice could draw images of decent quality. With text-to-image tools (like stable diffusion (Rombach et al., 2022)), anyone can create drawing images with a plain text description. Such a combination of user imagination power and AI execution power makes it possible to generate new types of images at an unprecedented speed. Beyond image generation, AIGC tasks also facilitate generating other types of content.

Another change AIGC brings is that the boundary between content consumer and creator becomes vague. In Web 2.0, Content generators and consumers are often different users. With AIGC in Web 3.0, however, data consumers are now able to become data creators, as they are able to use AI algorithms and technology to generate their own original content, and it allows them to have more control over the content they produce and consume, making them use their own data and AI technology to produce content that is tailored to their specific needs and interests. Overall, the shift towards AIGC has the potential to greatly transform the way data is consumed and produced, giving individuals and organizations more control and flexibility in the content they create and consume. In the following, we discuss why AIGC has become popular now.

2.2. Technology conditions

When it comes to AIGC technology, the first thing that comes into mind is often machine (deep) learning algorithm, while overlooking its two important conditions: data access and compute resources.

Advances in data access. Deep learning refers to the practice of training a model on data. The model performance heavily relies on the size of the training data. Typically, the model performance increases with more training samples. Taking image classification as an example, ImageNet (Deng et al., 2009) with more than 1 million images is a commonly used dataset for training the model and validating the performance. Generative AI often requires an even larger dataset, especially for challenging AIGC tasks like text-to-image. For example, approximately 250M images were used for training DALL-E (Ramesh et al., 2021). DALL-E 2 (Ramesh et al., 2022), on the other hand, used approximately 650M images. ChatGPT was built on top of GPT3 (Brown et al., 2020) partly trained on CommonCrawl dataset, which has 45TB of compressed plaintext before filtering and 570GB after filtering. Other datasets like WebText2, Books1/2, and Wikipedia are involved in the training of GPT3. Accessing such a huge dataset becomes possible mainly due to the Internet.

Advances in computing resources. Another important factor contributing to this development of AIGC is advanced in computing resources. Early AI algorithm was run on CPU, which cannot meet the need of training large deep learning models. For example, AlexNet (Krizhevsky et al., 2012) was the first model trained on full ImageNet and the training was done on Graphics Processing Units (GPUs). GPUs were originally designed for rendering graphics in video games but have become increasingly common in deep learning. GPUs are highly parallelized and can perform matrix operations much faster than CPUs. Nvidia is a leading company in manufacturing GPUs. The computing capability of its CUDA has improved from the first CUDA-capable GPU (GeForce 8800) in 2006 to the recent GPU (Hopper) with hundreds of times more computing power. The price of GPUs can range from a few hundred dollars to several thousand dollars, depending on the number of cores and memory. Tensor Processing Units (TPUs) are specialized processors designed by Google specifically for accelerating neural network training. TPUs are available on the Google Cloud Platform, and the pricing varies depending on usage and configuration. Overall, the price of computing resources is on the trend of becoming more affordable.

Fundamental techniques behind AIGC

In this work, we perceive AIGC as a set of tasks or applications that generates content with AI methods. Before introducing AIGC, we first visit the fundamental techniques behind AIGC, which fall in the scope of generative AI at the technical level. Here, we summarize the fundamental techniques by roughly dividing them into two classes: Generative techniques and Creation techniques. Specifically, Creation techniques refer to the techniques that are able to generate various contents, e.g., GAN and diffusion model. Meanwhile, General techniques cannot generate content directly but are essential for the development of AIGC, e.g., the Transformer architecture. In this section, we provide a brief summary of the required techniques for AIGC.

After the phenomenal success of AlexNet (Krizhevsky et al., 2012), there is a surging interest in deep learning, which somewhat becomes a synonym for AI. In contrast to traditional rule-based algorithms, deep learning is a data-driven method that optimizes the model parameters with a stochastic gradient. The success of deep learning in obtaining a superior feature representation depends on better backbone architecture and more data, which greatly accelerates the development of AIGC.

As two mainstream fields in deep learning, the research on natural language processing (NLP) and computer vision (CV) have significantly improved the backbone architectures and inspired various applications of improved backbones in other fields, e.g., the speech area. In the NLP field, Transformer (Vaswani et al., 2017) has replaced recurrent neural networks (RNN) (Mikolov et al., 2010; Medsker and Jain, 2001) to be the de-facto standard backbone. In the CV area, vision Transformer (ViT) (Dosovitskiy et al., 2020) has also shown its power besides the traditional convolutional neural networks (CNN). Here, we will briefly introduce how these mainstream backbones work and their representative variants.

RNN architecture. RNN is mainly adopted for handling data with time sequences, like language or audio. A vanilla RNN has three layers: input, hidden, and output. The information flow in RNN is in two directions. The first direction is from the input to the hidden layer and then to the output. What captures the recurrent nature of RNN lies in its second information flow in the time direction. Except for the corresponding input, the current hidden state depends at time tt depends on the hidden state at time t−1t-1. This two-flow design well handles the sequence order but suffers from exploding or vanishing gradients when the sequence gets long. To mitigate long-term dependency, LSTM (Hochreiter and Schmidhuber, 1997) was introduced with a cell state that acts like a freeway to facilitate the information flow in the sequence direction. LSTM is one of the most popular methods for alleviating the gradient vanishing/exploding issue. With three types of gates, however, LSTM suffers from high complexity and a higher memory requirement. Gated Recurrent Unit (GRU) (Chung et al., 2014) simplifies LSTM by merging its cell and hidden states and replacing the forget and input gates with a so-called update state. Unitary RNN (Arjovsky et al., 2016) handles the gradient issue by implementing unitary matrices. Gated Orthogonal Recurrent Unit (Jing et al., 2019a) leverages the merits of both gate and unitary matrices. Bidirectional RNN (Schuster and Paliwal, 1997) improves vanilla RNN by capturing both past and future information in the cell, i.e., the state at time tt is calculated based on both time t−1t-1 and t+1t+1. Depending on the tasks, RNN can have various architectures with a different number of inputs and outputs: one-to-one, many-to-one, one-to-many, and many-to-many. The many-to-many can be used in machine translation and is also called the sequence-to-sequence (seq2seq) model (Sutskever et al., 2014). Attention was introduced in (Bahdanau et al., 2014) to make the model decoder see every encoder token and automatically decide the weights on them based on their importance.

Transformer. Different from Seq2seq with attention (Bahdanau et al., 2014; Luong et al., 2015; Parikh et al., 2016), a new variant of architecture discards the seq-2seq architecture and claims that attention is all you need (Vaswani et al., 2017). Such attention is called self-attention, and the proposed architecture is termed Transformer (Vaswani et al., 2017) (see Figure 5). A standard Transformer consists of an encoder and a decoder and is developed based on residual connection (He et al., 2016) and layer normalization (Ba et al., 2016). Except for the Add &\& Norm module, the Transformer has two core components: multi-head attention and feed-forward neural network (a.k.a. MLP). The attention module adopts a multi-head design with the self-attention in the form of scaled dot-product defined as:

Unlike RNNs, which build positional information by sequentially inputting sentence information, Transformer obtains powerful modeling capabilities by constructing global dependencies but also loses information with positional bias. Therefore, positional encoding is needed to enable the model to sense the positional information of the input signal. There are two types of positional encoding. Fixed position coding is represented by sinusoids and cosines of different frequencies. The learnable position encoding is composed of a set of learnable parameters. Transformer has become the de-facto standard method in NLP tasks.

CNN architecture. After introducing RNN and Transformer in NLP field, we start to visit two mainstream backbones in CV area, i.e., CNN and ViT. CNNs have become a standard backbone in the field of computer vision. The core of CNN lies in its convolution layer. The convolution kernel (also known as filter) in the convolution layer is a set of shared weight parameters for operating on images, which is inspired by the biological visual cortex cells. The convolution kernel slides on the image and performs correlation operations with the pixel values on the image, finally obtaining the feature map and realizing the feature extraction of the image. GoogleNet(Szegedy et al., 2015), with its Inception module allowing multiple convolutional filter sizes to be chosen in each block, increased the diversity of convolutional kernels, thus the performance of CNN was improved. ResNet(He et al., 2016) was a milestone for CNNs, introducing residual connections that stabilized training and enabled the models to achieve better performance through deeper modeling. After that, it became part of the binding in CNNs. In order to expand the work of ResNet, DenseNet(Huang et al., 2017) establishes dense connections between all the previous layers and the subsequent layers, thus enabling the model to have better modeling ability. EfficientNet(Tan and Le, 2019) uses a scaling method which uses a set of fixed scaling coefficients to uniformly scale the width, depth, and resolution of the convolutional neural network architecture, thus making the model more efficient.

ViT architecture. Inspired by the success of Transformer in NLP, numerous works have tried to apply Transformer to the field of CV with ViT(Dosovitskiy et al., 2020) (see Figure 6), being the first of its kind. ViT first flattens the image into a sequence of 2D patches and inserts a class token at the beginning of the sequence to extract classification information. After the embedding position encoding, the token embeddings are fed into a standard Transformer. This simple and effective implementation of ViT makes it highly scalable. Swin (Liu et al., 2021b) efficiently deals with image classification and dense recognition tasks by constructing hierarchical feature maps by merging image blocks at a deeper level, and due to its computation of self-attention only within each local window, it reduces computational complexity. DeiT(Touvron et al., 2020) uses the teacher-student strategy for training, reducing the dependence of Transformer models on large data, by introducing distillation tokens. CaiT(Touvron et al., 2021) introduces class attention to effectively increase the depth of the model. T2T(Yuan et al., 2021b) effectively localizes the model by Token Fusion and introduces hierarchical deep and narrow structures through the prior of CNNs by recursively aggregating adjacent Tokens into one Token. Through permutation equivariance, Transformers have liberated CNNs from their translation invariance, allowing for long-range dependencies and less inductive bias, making them more powerful modeling tools and better transferable to downstream tasks than CNNs. In the current paradigm of large models and large datasets, Transformers have gradually replaced CNNs as the mainstream model in the field of computer vision.

1.2. Self-supervised pretraining

Parallel to better backbone architecture, deep learning also benefits from self-supervised pertaining which can exploit a larger (unlabeled) training dataset. Here, we summarize the most relevant pretraining techniques to AIGC, and categorize them according to the training data type (e.g., language, vision, and joint pretraining).

Language pretraining. There are three major types of language pretraining methods. The first type pretrains an encoder with masking, for which the representative work is BERT (Devlin et al., 2018) (see Figure 7). Specifically, BERT predicts the masked language tokens from the unmasked tokens. There is a significant discrepancy between the mask-then-predict pertaining task and downstream tasks, therefore masked language modeling like BERT is rarely used for text generation without finetuning. By contrast, autoregressive language pretraining methods are suitable for few-shot or zero-shot text generation. GPT family (Radford et al., 2018, 2019; Brown et al., 2020) is the most popular one which adopts a decoder instead of an encoder. Specifically, GPT-1 (Radford et al., 2018) is the first of its kind with GPT-2 (Radford et al., 2019) and GPT-3 (Brown et al., 2020) further investigating the role of massive data and large model in the transfer capacity. Based on GPT-3, the unprecedented success of ChatGPT has attracted great attention recently. Moreover, a stream of language models adopts both an encoder and decoder as the original Transformer. BART (Lewis et al., 2019) perturbed the input with various types of noise and predicted the original clean input, like a denoising autoencoder. MASS (Song et al., 2019) and PropheNet (Qi et al., 2020) follow BERT to take a masked sequence as the input of the encoder with the decoder predicting the masked tokens in an autoregressive manner. T5 (Raffel et al., 2020) replaces the masked tokens with some random tokens.

Visual pretraining. To learn better representations of vision data during pretraining, self-supervised learning (SSL) has been widely applied, and we term it visual SSL. Visual SSL has undergone three stages. Early works focused on designing various pretext tasks like jigsaw puzzles (Noroozi and Favaro, 2016) or predicting rotation (Gidaris et al., 2018). Such pretraining yields better performance on the downstream task than training from scratch, which motivates contrastive learning methods (Chen et al., 2020; He et al., 2020a; Zhang et al., 2022a). Contrastive learning adopts joint embedding to minimize the representation distance between augmented images for learning augmentation-invariant representation. The representation in pure joint embedding can collapse to a constant regardless of the inputs, for which contrastive learning simultaneously maximizes the representation distance from negative samples. Negative-free joint-embedding methods have also been investigated in SimSiam (Chen and He, 2021) and BYOL (Grill et al., 2020). How SimSiam works without negative samples have been investigated in (Zhang et al., 2022c). Inspired by the success of BERT in NLP for pertaining, BEiT (Bao et al., 2022) applied masking modeling in vision and its success relies on a pre-trained VAE to obtain the visual token. Masked autoencoder (MAE) (He et al., 2022a) (see Figure 8) simplifies it to an end-to-end denoising framework by predicting the masked patches from the unmasked patches. Outperforming contrastive learning and negative-free joint-embedding methods, MAE has become a new variant of the visual SSL framework. Interested readers can refer (Zhang et al., 2022b) for more details.

Joint pretraining. With large datasets of image-text pairs collected from the Internet, multimodal learning (Baltrušaitis et al., 2018; Xu et al., 2022) has made unprecedented progress to learn data representations, at the front of which is cross-modal matching (Gan et al., 2022). Contrastive pretraining is widely used to match the image embedding and text encoding in the same representation space (Radford et al., 2021; Jia et al., 2021; Yuan et al., 2021a). CLIP (Radford et al., 2021) (see Figure 9 is a pioneering work in this direction and is used in numerous text-to-image models, such as DALL-E 2 (Ramesh et al., 2022), Upainting (Li et al., 2022e), DiffusionCLIP (Kim and Ye, 2021). ALIGN (Jia et al., 2021) extended CLIP with noisy text supervision so that the text-image dataset requires no cleaning and can be scaled to a much larger size (from 400M to 1.8B). Florence (Yuan et al., 2021a) further expands the cross-modal shared representation from coarse scene to dine object and from static images to dynamic videos, etc. Therefore, the learned shared representation is more universal and shows superior performance (Yuan et al., 2021a).

2. Creation techniques in AI

Deep generative models (DGMs) are a group of probabilistic models that use neural networks to generate samples. Early attempts at generative modeling focused on pre-training with an autoencoder (Rumelhart et al., 1985; Ballard, 1987; Hinton and Zemel, 1993). A variant of autoencoder with masking has emerged to become a dominant self-supervised learning framework, and interested readers are encouraged to check a survey on masked autoencoder (Zhang et al., 2022b). Unless specified, the use cases of deep generative models in this survey only consider generating new data. The generated data is typically high-dimensional, and therefore, predicting a label of a sample is not considered discriminative instead of generative modeling even though something like a label is also technically generated.

Numerous DGMs have emerged and can be categorized into two major groups: likelihood-based and energy-based. Likelihood-based probabilistic models, like autoregressive models (Graves, 2013) and flow models (Dinh et al., 2015), have a tractable likelihood which provides a straightforward method to optimize the model weights w.r.t. the log-likelihood of the observed (training) data. The likelihood is not fully tractable in variational autoencoders (VAEs) (Kingma and Welling, 2013), but a tractable lower bound can be optimized, thus VAE is also considered to lie in the likelihood-based group which specifies a normalized probability. By contrast, energy-based models (Grenander and Miller, 1994; Hinton, 2002) are featured by the unnormalized probability, a.k.a. energy function. Without the constraint on the tractability of the normalizing constant, energy-based models are more flexible in parameterizing but difficult to train (Song and Kingma, 2021). Notably, GAN and diffusion models are highly related to energy-based models even though are developed from different motivations. In the following, we present an introduction to each class of likelihood-based models, followed by how the energy-based models can be trained as well as the mechanism behind GAN and diffusion models.

Autoregressive models. Autoregressive models learn the joint distribution of sequential data and predict each variable in the sequence with previous time-step variables as inputs. As shown in Eq. 2, autoregressive models assumes that the joint distribution pθ(x)p_{\theta}(x) can be decomposed to a product of conditional distributions.

Although both rely on previous timesteps, autoregressive models differ from RNN architecture since the previous timesteps are given to the model as input instead of hidden states in RNN. In other words, autoregressive models can be seen as a feed-forward network that takes all the previous time-step variables as inputs. Early works model discrete data with different functions estimating the conditional distribution, e.g. logistic regression in Fully Visible Sigmoid Belief Network (FVSBN) (Gan et al., 2015) and one hidden layer neural networks in Neural Autoregressive Distribution Estimation (NADE) (Larochelle and Murray, 2011). The following research further extends to model the continuous variables (Uria et al., 2013, 2016). Autoregressive methods have been widely applied in multiple areas, including computer vision (PixelCNN (Van den Oord et al., 2016) and PixelCNN++ (Salimans et al., 2017)), audio generation (WaveNet (van den Oord et al., 2016)), natural language processing (Transformer (Vaswani et al., 2017)).

VAE. Autoencoders are a family of models that first map the input to a low-dimension latent layer with an encoder and then reconstruct the input with a decoder. The entire encoder-decoder process aims to learn the underlying data patterns and generate unseen samples (Oussidi and Elhassouny, 2018). Variational autoencoder (VAE) (Kingma and Welling, 2013) is an autoencoder that learns the data distribution p(x)p(x) from latent space z, i.e., p(x)=p(x∣z)p(z)p(x)=p(x|z)p(z), where p(x∣z)p(x|z) is learned by the decoder. In order to obtain p(z)p(z), VAE (Kingma and Welling, 2013) adopts Bayes’ theorem and approximates the posterior distribution p(z∣x)p(z|x) by the encoder. The VAE model is optimized toward a likelihood goal with regularizer (Altosaar, 2016).

2.2. Energy-based models

With a tractable likelihood, autoregressive models and flow models allow a straightforward optimization of the parameters w.r.t. the log-likelihood of the data. This forces the model to be constrained in a certain form. For example, the autoregressive model needs to be factorized as a product of conditional probabilities, and the flow model must adopt invertible transformation.

Energy-baed models specify probability up to an unknown normalizing constant, therefore, they are also known as non-normalzied probabilistic models. Without losing generality by assuming the energy-based model is over a single variable x\bm{x}, we denote its energy as Eθ(x)E_{\theta}(\bm{x}). Its probability density is then calculated as

where zθz_{\theta} is the so-called normalizing constant and defined as zθ=∫exp⁡(−Eθ(x)) dxz_{\theta}=\int\exp(-E_{\theta}(\bm{x}))~{}\text{d}\bm{x}. zθz_{\theta} is an intractable integral, making optimizing energy-based models a challenging task.

MCMC and NCE. Early attempts at optimizing energy-based models opt to estimate the gradient of the log-likelihood with Markov chain Monte Carlo (MCMC) approaches, which require a cumbersome drawing of random samples. Therefore, some works aim to improve the efficiency of MCMC a representative work Langevin MCMC (Parisi, 1981; Grenander and Miller, 1994). Nonetheless, performing MCMCM to obtain requires large computation and contrastive divergence (CD) (Hinton, 2002) is a popular method to reduce the computation via approximation with various variants: persistent CD (Tieleman, 2008), mean field CD (Welling and Hinton, 2002), and multi-grid CD (Gao et al., 2018). Another line of work optimizes energy-based models via notice contrastive estimation (NCE) (Gutmann and Hyvärinen, 2010), which contrasts the probabilistic model with another noise distribution. Specifically, it optimizes the following loss:

Score matching. For optimizing energy-based models, another popular MCMC-free method minimizes the derivatives of log probability density between the model and the observed data. The first-order of a log probability density function is called score of the distribution (s(x)=∇xlogp(x)s(\bm{x})=\nabla_{\bm{x}}\text{log}p(\bm{x})), therefore, this method is often termed score matching. Unfortunately, the data score function sd(x)s_{d}(\bm{x}) is unavailable. Various attempts (Vincent, 2011; Saremi et al., 2018; Song and Ermon, 2019; Pang et al., 2020; Shi et al., 2018; Song et al., 2020) have been made to mitigate this issue, with a representative method called denoising score matching (Vincent, 2011). Denoising score matching approximates the score of data with noisy samples. The model takes a noisy sample as the input and predicts its noise. Therefore, it can be used for sampling clean samples from noise by iterative removing the noise (Saremi et al., 2018; Song and Ermon, 2019).

2.3. Two star-models: from GAN to diffusion model

When it comes to deep generative models, what first comes to your mind? The answer depends on your background, however, GAN is definitely one of the most mentioned models. GAN stands for generative adversarial network (Goodfellow et al., 2014) which was first proposed by Ian J. Goodfellow and his team in 2014 and rated as “the most interesting idea in the last 10 years in machine learning” by Yann Lecun in 2016. As the pioneering work to generate images of reasonably high quality, GAN has been widely regarded as a de facto standard model for the challenging task of image synthesis. This long-time dominance has been recently challenged by a new family of deep generative models termed diffusion models (Ho et al., 2020). The overwhelming success of diffusion models starts from image synthesis but extends to other modalities, like video, audio, text, graph, etc. Considering their dominant influence in the development of generative AI, we first summarize GAN and diffusion models before introducing other families of deep generative models.

GAN. The architecture of GAN is shown in Figure 10. GAN is featured by its two network components: a discriminator (D\mathcal{D}) and a generator (G\mathcal{G}). D\mathcal{D} distinguishes real images from those generated by G\mathcal{G}, while G\mathcal{G} aims to fool D\mathcal{D}. Given a latent variable z∼pz\bm{z}\sim p_{\bm{z}}, the output of G\mathcal{G} is G(z)\mathcal{G}(\bm{z}) constituting a probability distribution pgp_{\bm{g}}. The goal of GAN is to make pgp_{\bm{g}} approximate the observed data distribution pdatap_{\bm{data}}. This objective is achieved through adversarial learning, which can be interpreted as a min-max game (Schmidhuber, 1990):

where D\mathcal{D} is trained to maximize the probability of assigning correct labels to real images and generated ones, and is used to guide the optimization of G\mathcal{G} towards generating more real images. GANs have the weakness of potentially unstable training and less diversity in generation due to their adversarial training nature. The basic difference between GANs and autoregressive models is that GANs learn implicit data distribution, whereas the latter learns an explicit distribution governed by a prior imposed by model structure.

Diffusion model. The use of diffusion models, a special form of hierarchical VAEs, has seen explosive growth in the past few years (Ulhaq et al., 2022; Croitoru et al., 2022; Cao et al., 2022; Li et al., 2022d; Pascual et al., 2022). Diffusion models (Figure 11) are also known as denoising diffusion probabilistic models (DDPMs) or score-based generative models that generate new data similar to the data on which they are trained (Ho et al., 2020). Inspired by non-equilibrium thermodynamics, DDPMs can be defined as a parameterized Markov chain of diffusion steps to slowly add random noise to the training data and learn to reverse the diffusion process to construct desired data samples from the pure noise.

In the forward diffusion process, DDPM destroys the training data through the successive addition of Gaussian noise. Given a data distribution x0∼q(x0)\mathbf{x}_{0}\sim q(\mathbf{x}_{0}), DDPM maps the training data to noise by gradually perturbing the input data. This is formally achieved by a simple stochastic process that starts from a data sample and iteratively generates noisier samples xT\mathbf{x}_{T} with q(xt∣xt−1)q(\mathbf{x}_{t}\mid\mathbf{x}_{t-1}), using a simple Gaussian diffusion kernel:

where TT and βt\beta_{t} are the diffusion steps and hyper-parameters, respectively. We only discuss the case of Gaussian noise as transition kernels for simplicity, indicated as N\mathcal{N} in Eq. 7. With αt:=1−βt\alpha_{t}:=1-\beta_{t} and αˉt:=∏s=0tαs\bar{\alpha}_{t}:=\prod_{s=0}^{t}\alpha_{s}, we can obtain noised image at arbitrary step tt as follows:

During the reverse denoising process, DDPM is learning to recover the data by reversing the noising process i.e., it undoes the forward diffusion by performing the iterative denoising. This process represents data synthesis and DDPM is trained to generate data by converting random noise into real data. It is also formally defined as a stochastic process, which iteratively denoises the input data starting from pθ(T)p_{\theta}(T) and generates pθ(x0)p_{\theta}(x_{0}) which can follow the true data distribution q(x0)q(x_{0}). Therefore, the optimization objective of the model is as follows:

Both the forward and reverse processes of DDPMs often use thousands of steps for gradual noise injection and during generation for denoising.

AIGC task: text generation

NLP studies natural language with two fundamental tasks: understanding and generation. These two tasks are not exclusively separate because the generation of an appropriate text often depends on the understanding of some text inputs. For example, language models often transform a sequence of text into another, which constitutes the core task of text generation, including machine translation, text summarization, and dialogue systems. Beyond this, text generation evolves in two directions: controllability and multi-modality. The first direction aims to make the generated content

The main task of the dialogue system (chatbots) is to provide better communication between humans and machines (Ni et al., 2022; Deriu et al., 2021). According to whether the task is specified in the applications, dialogue system can be divided into two categories : (1) task-oriented dialogue systems (TOD) (Zhang et al., 2020b; Peng et al., 2020; Yang et al., 2021) and (2) open-domain dialogue systems (OOD) (Zhou et al., 2020a; Zhang et al., 2019b; Adiwardana et al., 2020). Specifically, the task-oriented dialogue systems focus on task completion and solve specific problems (e.g., restaurant reservations and ticket booking) (Zhang et al., 2020b). Meanwhile, open-domain dialogue systems are often data-driven and aim to chat with humans without task or domain restrictions (Zhang et al., 2020b; Ritter et al., 2011).

Task-oriented systems. Task-oriented dialogue systems can be divided into modular and end-to-end systems. The modular methods include four main parts: natural language understanding (NLU) (Singla et al., 2020; Su et al., 2019), dialogue state tracking (DST) (Shan et al., 2020; Wang et al., 2020b), dialogue policy learning (DPL) (Huang et al., 2020; Xu et al., 2020), and natural language generation (NLG) (Baheti et al., 2020; Elder et al., 2020). After encoding the user inputs into semantic slots with NLU, DST, and DPL decide the next action that is then converted to natural language by NLG as the final response. These four modules aim to generate responses in a controllable way and can be optimized individually. However, some modules may not be differentiable, and the improvement of a single module may not lead to the improvement of the whole system (Zhang et al., 2020b). To solve these problems, end-to-end methods either achieve an end-to-end training pipeline by making each module differentiable (Ham et al., 2020; Hosseini-Asl et al., 2020), or use a single end-to-end module in the system (Zhang et al., 2020a; Yang et al., 2020b). There still exist several challenges for both modular and end-to-end systems, including how to improve tracking efficiency for DST (Kim et al., 2019; Ouyang et al., 2020) and how to increase the response quality of end-to-end system with limited data (He et al., 2020b; Henderson et al., 2019; Mehri et al., 2019).

Open-domain systems. Open-domain systems aim to chat with users without task and domain restrictions (Ritter et al., 2011; Zhang et al., 2020b), and can be categorized into three types: retrieval-based systems, generative systems, and ensemble systems (Zhang et al., 2020b). Specifically, retrieval-based systems always find an existing response from a response corpus, while generative systems can generate responses that may not appear in the training set. Ensemble systems combine retrieval-based and generative methods by either choosing the best response or refining the retrieval-based model with generative one (Zhang et al., 2020b; Zhu et al., 2018; Serban et al., 2017). Previous works improve the open-domain systems from multiple aspects, including dialogue context modeling (Feng et al., 2020; Jia et al., 2020; Lin et al., 2020; Mehri et al., 2019), improving the response coherence (Xu et al., 2020; Gao et al., 2020; Akama et al., 2020; Lison and Bibauw, 2017) and diversity (Qiu et al., 2019; Bao et al., 2019; Ko et al., 2020; Su et al., 2020). Most recently, ChatGPT (see Figure 12) has achieved unprecedented success and also falls into the scope of open-domain dialogue systems. Apart from answering various questions, ChatGPT can also be used for paper writing, code debugging, table generation, and to name but a few.

1.2. Machine translation

As the term suggests, machine translation automatically translates the text from one language to another (Hutchins, 1986; Yang et al., 2020a) (see Figure 13). With deep learning replacing rule-based (Forcada et al., 2011) and statistical (Koehn et al., 2003, 2007) methods, neural machine translation (NMT) requires minimum linguistic expertise (Song and Croft, 1999; Wallach, 2006) and has become a mainstream approach featured by its higher capacity in capturing long dependency in the sentence (Cho et al., 2014). The success of neural machine learning can be mainly attributed to language models (Bengio et al., 2000), which predicts the probability of a word conditioned on previous ones. Seq2seq (Sutskever et al., 2014) is a pioneering work to apply encoder-decoder RNN structure (Kalchbrenner and Blunsom, 2013) to machine translation. When the sentence gets long, the performance of Seq2seq (Sutskever et al., 2014) deteriorates, for which an attention mechanism was proposed in (Bahdanau et al., 2014) to help translate the long sentence with additional word alignment. With increasing attention, in 2006, Google’s NMT system helped reduce the translation effort of humans by around 60%60\% compared to Google’s phrase-based production system, which bridges the gap between Human and machine translation (Wu et al., 2016). CNN-based architectures have also been investigated for NMT with numerous attempts (Kaiser and Bengio, 2016; Kalchbrenner et al., 2016), but fail to achieve comparable performance as the RNN boosted by attention (Bahdanau et al., 2014). Convolutional Seq2seq (Gehring et al., 2017) makes CNN compatible with the attention mechanism, showing CNN can achieve comparable or even better performance than RNN. However, this improvement was later outperformed by another architecture termed Transformer (Vaswani et al., 2017). With RNN or Transformer as the architecture, NMT often utilizes autoregressive generative model, where a greedy search only considers the word with the highest probability for predicting the next work during inference.

A trend for NMT is to achieve satisfactory performance in low-resource setup, where the model is trained with limited bilingual corpus (Wang et al., 2021b). One way to mitigate this data scarcity is to utilize auxiliary languages, like multilingual training with other language pairs (Zoph and Knight, 2016; Shatz, 2017; Johnson et al., 2017) or pivot translation with English as the middle pivot language (Ren et al., 2018; Cheng, 2019). Another popular approach is to utilize pre-trained language models, like BERT (Devlin et al., 2018) or GPT (Radford et al., 2018). For example, it is shown in (Rothe et al., 2020) that initializing the model weights with BERT (Devlin et al., 2018) or RoBERTa (Liu et al., 2019a) significantly improves the English-German translation performance. Without the need for fine-tuning, GPT-family models (Radford et al., 2018, 2019; Brown et al., 2020) also show competitive performance. Most recently, ChatGPT has shown its power in machine translation, performing competitively with commercial products (e.g., Google translate) (Jiao et al., 2023).

2. Multimodal text generation

Image-to-text, also known as image captioning, refers to describing a given image’s content in natural language (see Figure 14). A seminal work in this area is Neural Image Caption (NIC) (Vinyals et al., 2015), which employs CNN as an encoder to extract high-level representations of input images and then feed these representations into an RNN decoder to generate image descriptions. This two-step encoder-decoder architecture has been widely applied in later works on image captioning, and we term them as visual encoding (Stefanini et al., 2021) and language decoding, respectively. Here, we first revisit the history and recent trends of both stages in image captioning.

Visual encoding. Extracting an effective representation of images is the main task of visual encoding module. Start from NIC (Vinyals et al., 2015) with GoogleNet (Szegedy et al., 2015) extracting the global feature of input image, multiple works adopt various CNN backbones as the encoder, including AlexNet (Krizhevsky et al., 2012) in (Karpathy and Fei-Fei, 2015) and VGG network (Simonyan and Zisserman, 2015) in (Mao et al., 2014; Donahue et al., 2015). However, it is hard for a language model to generate fine-grained captions with global visual features. Following works introduce attention mechanism for fine-grained visual features, including attention over different grids of CNN features (Xu et al., 2015; Lu et al., 2017; Wang et al., 2017; Chen et al., 2018) or over different visual regions (Anderson et al., 2018; Ke et al., 2019; Zha et al., 2019). Another branch of work (Yang et al., 2019a; Zhao et al., 2021) adopts graph neural networks to encode the semantic and spatial relationships between different regions. However, the human-defined graph structures may limit the interactions among elements (Stefanini et al., 2021), which can be mitigated by the self-attention methods (Yang et al., 2019b; Li et al., 2019c; Zhang et al., 2021b) (including ViT (Liu et al., 2021a)) that connects all the elements.

Language decoding. In image captioning, a language decoder generates captions by predicting the probability of a given word sequence (Stefanini et al., 2021). Inspired by the breakthroughs in the NLP area, the backbones of language decoders evolve from RNN (Vinyals et al., 2015; Lu et al., 2017; Ke et al., 2019; Wang et al., 2020a) to Transformer (Li et al., 2019c; Herdade et al., 2019; Guo et al., 2020), achieving significant performance improvement. Beyond the visual encoder-language decoder architecture, a branch of work adopts BERT-like architecture that fuses the image and captions in the early stage of a single model (Li et al., 2020d; Zhou et al., 2020b; Zhang et al., 2021a). For example, (Zhou et al., 2020b) adopts a single encoder to learn a shared space for image and text, which is first pre-tained on large image-text corpus and finetuned, specifically for image captioning tasks.

2.2. Speech-to-Text

Speech-to-text generation, also known as automatic speech recognition (ASR), is the process of converting spoken language, specifically a speech signal, into a corresponding text (Reddy, 1976; Indurkhya and Damerau, 2010) (see Figure 15). With many potential applications such as voice dialing, computer-assisted language learning, caption generation, and virtual assistants like Alexa and Siri, ASR has been an exciting field of research (Raut and Deoghare, 2016; Karpagavalli and Chandra, 2016; Malik et al., 2021) since the 1950s, and evolved from hidden Markov models (HMM) (Levinson et al., 1983; Juang and Rabiner, 1991) to DNN-based systems (Dahl et al., 2011; Hinton et al., 2012; Graves et al., 2013; Nassif et al., 2019; Wu et al., 2020).

Various research topics and challenges. Previous works improved ASR systems in various aspects. Multiple works discuss different feature extraction methods for speech signals (Malik et al., 2021), including temporal features (e.g., discrete wavelet transform (Milone and Di Persia, 2008; Tang, 2009)) and spectral features such as the most commonly used mel-frequency cepstral coefficients (MFCC) (Tóth, 2011; Collobert et al., 2016; Chiu et al., 2018). Another branch of work improves the system pipeline (Roger et al., 2022) from multi-model (Lüscher et al., 2019) to end-to-end ones (Hori et al., 2018; Li et al., 2019a; Nakatani, 2019; Wang et al., 2019; Li et al., 2020c). Specifically, a multi-model system (Lüscher et al., 2019; Malik et al., 2021) first learns an acoustic model (e.g., a phoneme classifier that maps the features to phonemes) and then a language model for the word outputs (Roger et al., 2022). On the other hand, end-to-end models directly predict the transcriptions from the audio input (Hori et al., 2018; Li et al., 2019a; Nakatani, 2019; Wang et al., 2019; Li et al., 2020c). Although end-to-end models achieve impressive performance in various languages and dialects, many challenges still exist. First, their applications for under-resourced speech tasks remain challenging as it is costly and time-consuming to acquire vast amounts of annotated training data (Fendji et al., 2022; Roger et al., 2022). Second, these systems may struggle to handle speech with specialized out-of-vocabulary words and may perform well on the training data but may not generalize well to new or unseen data (Qin, 2013; Fendji et al., 2022). Moreover, biases in the training data can also affect the performance of supervised ASR systems, leading to poor accuracy on certain groups of people or speech styles (Benzeghiba et al., 2007).

Under-resourced speech tasks. Researchers work on new technologies to overcome challenges in ASR systems, among which we mainly discuss the under-resourced speech problem that lacks data for impaired speech (Roger et al., 2022). A branch of work (Pascual et al., 2019; Ravanelli et al., 2020) adopts multi-task learning to optimize a shared encoder for different tasks. Meanwhile, self-supervised ASR systems have recently become an active area of research without relying on a large number of labeled samples. Specifically, self-supervised ASR systems first pre-train a model on huge volumes of unlabeled speech data, then fine-tune it on a smaller set of labeled data to facilitate the efficiency of ASR systems. It can be applied for low-resource languages, handling different speaking styles or noise conditions, and transcribing multiple languages (Liu et al., 2022a; Yadav and Sitaram, 2022; Conneau et al., 2020; Baevski et al., 2019).

AIGC task: image generation

Similar to text generation, the task of image synthesis can also be categorized into different classes based on its input control. Since the output is images, a straightforward type of control is images. Image-type control induces numerous tasks, like super-resolution, deblur, editing, translation, etc. A limitation of image-type control is the lack of flexibility. By contrast, text-guided control enables the generation of any image content with any style at the free will of humans. Text-to-image falls into the category of cross-modal generation, since the input text is a different modality from the output image.

Image restoration solves a typical inverse problem that restores clean images from their corresponding degraded versions, with examples shown in Figure 16. Such an inverse problem is non-trivial with its ill-posed nature because there are infinite possible mappings from the degraded image to the clean one. There are two sources of degradation: missing information from the original image and adding something undesirable to the clean image. The former type of degradation includes capturing a photo with a low resolution and thus losing some detailed information, cropping a certain region, and transforming a colorful image to its gray form. Restoration tasks recover them in order are image super-resolution, inpainting, and colorization, respectively. Another class of restoration tasks aims to remove undesirable perturbations, like denoise, derain, dehaze, deblur, etc. Early restoration techniques primarily use mathematical and statistical modeling to remove image degradations, including spatial filters for denoising (Gonzalez, 2006; Shrestha, 2014; Zhang et al., 2013), kernel estimation for deblurring (Xu and Jia, 2010; Xu et al., 2016). Lately, deep learning-based methods (Jain and Seung, 2008; Xie et al., 2012; Xu et al., 2014; Liang and Liu, 2015; Dong et al., 2014; Cheng et al., 2015; Cai and Wei, 2020; Liu et al., 2021c) have become predominant in image restoration tasks due to their versatility and superior visual quality over their traditional counterparts. CNN is widely used as the building block in image restoration (Wang et al., 2015; Sun et al., 2015; Dong et al., 2015; Varga and Szirányi, 2016), while recent works explore more powerful transformer architecture and achieve impressive performance in various tasks, such as image super-resolution (Liang et al., 2021), colorization (Kumar et al., 2021), and inpainting (Li et al., 2022a). There are also works that combine the strength of CNNs and Transformers together (Fang et al., 2022; Zhao et al., 2022b, a).

Generative methods for restoration. Typical image restoration models learn a mapping between the source (degraded) and target (clean) images with a reconstruction loss. Depending on the task, training data pairs can be generated by degrading clean images with various perturbations, including resolution downsampling and grayscale transformation. To keep more high-frequency details and create more realistic images, generative models are widely used for restoration, such as GAN in super-resolution (Ledig et al., 2017; Wang et al., 2018a; Zhang et al., 2019a) and inpainting (Nazeri et al., 2019; Cai and Wei, 2020; Liu et al., 2021c). However, GAN-based models typically suffer from a complex training process and mode collapse. These drawbacks and the massive popularity of DMs led numerous recent works to adopt DMs for image restoration tasks (Li et al., 2022f; Saharia et al., 2022c; Kawar et al., 2022; Saharia et al., 2022a; Lugmayr et al., 2022; Ren et al., 2022). Generative approaches like GAN and DM can also produce multiple variations of clean output from a single degraded image.

From single-task to multi-task. A majority of existing restoration approaches train separate models for different forms of image degradation. This limits their effectiveness in practical use cases where the images are corrupted by a combination of degradations. To address this, several studies (Shin et al., 2022; Kim et al., 2020; Zhou et al., 2022; Ahn et al., 2017) introduce multi-distortion datasets that combine various forms of degradation with different intensities. Some studies (Yu et al., 2018; Liu et al., 2019b; Kim et al., 2020; Yuan et al., 2018) propose restoration models in which different sub-networks are responsible for different degradations. Another line of work (Shin et al., 2022; Li et al., 2022b; Zhou et al., 2022; Li et al., 2020a; Suganuma et al., 2019) relies on attention modules or a guiding sub-network to assist the restoration network through different degradations, allowing a single network to handle multiple degradations.

1.2. Image editing

In contrast to image restoration for enhancing image quality, image editing refers to modifying an image to meet a certain need like style transfer (see Figure 17). Technically, some image restoration tasks like colorization might also be perceived as image editing by perceiving adding color as the desired need. Modern cameras often have basic editing features such as sharpness adjustments (Zhang et al., 2017a), automatic cropping (Zhang et al., 2005), red eye removal (Smolka et al., 2003), etc. However, in AIGC, we are more interested in advanced image editing tasks that change the image semantics in various forms, such as content, style, object attributes, etc.

A family of image editing targets to modify the attributes (like age) of the main object (like a face) in the image. A typical use case is facial attribute editing which can change the hairstyle, age, or even gender. Based on a pre-trained CNN encoder, a line of pioneering works adopt optimization-based approaches (Li et al., 2016a; Upchurch et al., 2017), which is time-consuming due to its iterative nature. Another line of works adopts learning-based approaches to directly generate the image, with a trend from single attribute (Li et al., 2016b; Shen and Liu, 2017) to multiple ones (Xiao et al., 2017; Kim et al., 2017; He et al., 2019). A drawback of most aforementioned methods is the dependence on annotated labels for attributes, therefore, unsupervised learning has been introduced to disentangle different attributes (Shen and Zhou, 2021; Cherepkov et al., 2021).

Another family of image editing changes the semantics by combining two images. For example, image morphing (Jing et al., 2019b) interpolates the content of two images, while style transfer (Gatys et al., 2016) yields a new image with the content of one image and the style of the other. A naive method for image morphing is to perform interpolation in the pixel space, which causes obvious artifacts. By contrast, interpolating in the latent space can consider the view change and generate a smooth image. The latent space for those two images can be obtained via GAN inversion method (Xia et al., 2022). Numerous works (Zhu et al., 2016; Abdal et al., 2019; Zhu et al., 2020; Xu et al., 2021) have explored the latent place of a pre-trained GAN for image morphing. For the task of style transfer, a specific style-based variant of GAN termed StyleGAN (Karras et al., 2019) is a popular choice. From the earlier layers to the latter ones, StyleGAN controls the attributes from coarser-grained (like structure) to finer-grained ones (like texture). Therefore, StyleGAN can be used for style transfer by mixing the earlier layer’s latent representation of the content image and the latter layer’s latent representation of the style image (Abdal et al., 2019; Viazovetskyi et al., 2020; Guan et al., 2020; Wei et al., 2022).

Compared with restoration tasks, various editing tasks enable a more flexible image generation. However, its diversity is still limited, which is alleviated by allowing other text as the input. More recently, image editing based on diffusion models has been widely discussed and achieved impressive results (Kim and Ye, 2021; Chandramouli and Gandikota, 2022; Hertz et al., 2022; Wallace et al., 2022). DiffusionCLIP (Kim and Ye, 2021) is a pioneering work that finetunes a pre-trained diffusion model to align the target image and text. By contrast, LDEdit (Chandramouli and Gandikota, 2022) avoids finetuning based on LDM (Rombach et al., 2022). A branch of works discusses the mask problem in image editing, including how to connect a manually designed masked region and background seamlessly (Avrahami et al., 2022c, c, a; Ackermann and Li, 2022). On the other hand, DiffEdit (Couairon et al., 2022) proposes to predict the mask automatically that indicates which part to be edited. There are also works editing 3D objects based on diffusion models and text guidance (Li et al., 2022g; Kim and Chun, 2022; Chan et al., 2022).

2. Multimodal image generation

Text-to-image (T2I) task aims to generate images from textual descriptions (see Figure 18.), and can be traced back to image generation from tags or attributes (Srivastava and Salakhutdinov, 2012; Yan et al., 2016). AlignDRAW (Mansimov et al., 2016) is a pioneering work to generate images from natural language, and it is impressive that AlignDRAW (Mansimov et al., 2016) can generate images from novel text like ‘a stop sign is flying in blue skies’. More recently, advances in text-to-image area can be categorized into three branches, including GAN-based methods, autoregressive methods, and diffusion-based methods.

GAN-based methods. The limitation of AlignDRAW (Mansimov et al., 2016) is that the generated images are unrealistic and require an additional GAN for post-processing. Based on a deep convolutional generative adversarial network (DC-GAN) (Radford et al., 2015), (Reed et al., 2016) is the first end-to-end differential architecture from the character level to the pixel level. To generate high-resolution images while stabilizing the training process, StackGAN (Zhang et al., 2017b) and StackGAN++ (Zhang et al., 2018) propose a multi-stage mechanism that multiple generators produce images of different scales, and high-resolution image generation is conditioned on the low-resolution images. Moreover, AttnGAN (Xu et al., 2018) and Controlgan (Li et al., 2019b) adopt attention networks to obtain fine-grained control on the subregions according to relevant words.

Autoregressive methods. Inspired by the success of autoregressive Transformers (Vaswani et al., 2017), a branch of works generates images in an auto-regressive manner by mapping images to a sequence of tokens, among which DALL-E (Ramesh et al., 2021) is a pioneering work. Specifically, DALL-E (Ramesh et al., 2021) first converts the images to image tokens with a pre-trained discrete variational autoencoder (dVAE), then trains an auto-regressive Transformer to learn the joint distribution of text and image tokens. A concurrent work CogView (Ding et al., 2021) independently proposes the same idea with DALL-E (Ramesh et al., 2021) but achieves superior FID (Heusel et al., 2017) than DALL-E (Ramesh et al., 2021) on blurred MS COCO dataset. CogView2 (Ding et al., 2022) extends CogView (Ding et al., 2021) to various tasks, e.g., image captioning, by masking different tokens. Parti (Yu et al., 2022b) further improves the image quality by scaling the model size to 20 billion.

Diffusion-based methods. Diffusion model-based methods have achieved unprecedented success and attention recently, which can be categorized by either working on the pixel space directly (Nichol et al., 2022; Saharia et al., 2022b) or the latent space (Rombach et al., 2022; Ramesh et al., 2022). GLIDE (Nichol et al., 2022) outperforms DALL-E by extending class-conditional diffusion models to text-conditional settings, while Imagen (Saharia et al., 2022b) improves the image quality further with a pre-trained large language model (e.g., T5) capturing the text semantics. To reduce resource consumption of diffusion models in pixel space, Stable Diffusion (Rombach et al., 2022) first compresses the high-resolution images to a low-dimensional latent space, then trains the diffusion model in the latent space. This method is also known as Latent Diffusion Models (LDM) (Rombach et al., 2022). Different from Stable Diffusion (Rombach et al., 2022) that learns the latent space based on only images, DALL-E2 (Ramesh et al., 2022) applies diffusion model to learn a prior as alignment between image space and text space of CLIP. Other works also improve the model from multiple aspects, including introducing spatial control (Avrahami et al., 2022b; Voynov et al., 2022) and reference images (Blattmann et al., 2022; Sheynin et al., 2022).

2.2. Talking face

From the perspective of output, the task of talking face(Zhen et al., 2023) generates a series of image frames which are thus technically a video (see Figure 19). Different from general video generation (see Sec. 6.1), talking face requires an image face as an identity reference, and edits it based on the speech input. In this sense, talking face is more related to image editing. Moreover, talking face converts a speech clip to a corresponding face image, resembling speech recognition to convert a speech clip to a corresponding word text. With speech recognition recognized as a multimodal generation text task, this survey considers talking face as a multimodal image generation task. Driven by deep learning models, speech-to-head video synthesis models have attracted wide attention, which can be divided into 2D-based methods and 3D-based methods.

With 2D-based methods, talking face video synthesis mainly relies on landmarks, semantic maps, or similar representations. Landmarks are used as an intermediate layer from low-dimensional audio to high-dimensional video, as well as two decoders to decouple speech and speaker identity for generating video unaffected by speaker identity (Chung et al., 2017), which is also the first work to use deep generative models to create speech faces. In addition, image-to-image translation generation (Jamaludin et al., 2019) can also be used for lip synthesis, while the combination of separate audio-visual representations and neural networks can also be used to optimize synthesis (Song et al., 2018; Zhou et al., 2019a)

Another line of work is based on building a 3D model and controlling the motion process through rendering technology (Suwajanakorn et al., 2017; Kumar et al., 2017), with a drawback of high construction cost. Later, many generative talking face models based on 3DMM parameters (Karras et al., 2017; Cudeiro et al., 2019; Fried et al., 2019; Thies et al., 2020) were established, using models such as blendshape (Cudeiro et al., 2019), flame (Li et al., 2017), and 3D mesh (Richard et al., 2021), with audio as model input for content generation. At present, most methods are directly reconstructed from training videos. NeRF uses multi-layer perceptrons to simulate implicit representations, which can store 3D spatial coordinates and appearance information and are used for high-resolution scenes (Mildenhall et al., 2021; Müller et al., 2022; Li et al., 2022c). In addition, a pipeline and an end-to-end framework for unrestricted talking face video synthesis have also been proposed (Prajwal et al., 2020b; KR et al., 2019), taking any unidentified video and arbitrary speech as input.

AIGC task: beyond text and image

Compared with image generation, the progress of video generation lags behind largely because of the complexity of modeling higher-dimensional video data. Video generation involves not only generating pixels but also ensuring semantic coherence between different frames. Video generation works can be categorized into unguided and guided generation (e.g., text, images, video, and action classes), with text-guided age (see Figure LABEL:fig:_video_generation) receiving the most attention due to its high influence.

Unguided video generation. Early works on extending image generation from single frame to multiple frames are limited to creating monotonous yet regular content like sea waves. The generated dynamic textures (Wei and Levoy, 2000; Doretto et al., 2003) often have a spatially repetitive pattern with time-varying visualization. With the development of generative models, numerous works (Vondrick et al., 2016; Saito et al., 2017; Ohnishi et al., 2018; Tulyakov et al., 2018; Acharya et al., 2018; Clark et al., 2019b; Yushchenko et al., 2019) extend the exploration from naive dynamic textures to real video generation. Nonetheless, their success is limited to short videos for simple scenes with the availability of low-resolution datasets. More recent works (Clark et al., 2019a; Saito et al., 2020; Tian et al., 2021; Ho et al., 2022b) improve the video quality further, among which (Ho et al., 2022b) is regarded as a pioneering work of diffusion models.

Text-guided video generation. Compared to text-to-image models that can create almost photorealistic pictures, text-guided video generation is more challenging. Early works (Mittal et al., 2017; Pan et al., 2017; Marwah et al., 2017; Li et al., 2018; Gupta et al., 2018; Liu et al., 2019c) based on VAE or GAN concentrate on creating video in simple settings, such as digit bouncing, and human walking. Given the great success of the VQ-VAE model in text-guided image generation, some works (Wu et al., 2021; Hong et al., 2022) extend it to text-guided video generation, resulting in more realistic video scenes. To achieve high-quality video, (Ho et al., 2022b) first applies the diffusion model to text-guided video generation, which refreshes the benchmarks of evaluation. After that, Meta and Google propose Make-a-Video (Singer et al., 2022) and Imagen Video (Ho et al., 2022a) based on the diffusion model, respectively. Specifically, Make-a-Video extends a diffusion-based text-guided image generation model to video generation, which can speed up the generation and eliminate the need for paired text-video data in training. However, Make-a-Video requires a large-scale text-video dataset for fine-tuning, which results in a significant amount of computational resources. The latest Tune-a-Video (Wu et al., 2022) proposes one-shot video generation, driven by text guidance and image inputs, where a single text-video pair is used to train an open-domain generator.

2. 3D generation

The tremendous success of deep generative models on 2D images has prompted researchers to explore 3D data generation, which is actually a modeling of the real physical world. Different from the single format of 2D data, a 3D object can be represented by depth images, voxel grids(Wu et al., 2015), point clouds(Qi et al., 2017a, b), meshes(Hanocka et al., 2019) and neural fields(Mescheder et al., 2019), each of which has its advantages and disadvantages.

According to the type of input and guidance, 3D objects can be generated from text, images and 3D data. Although multiple methods (Jahan et al., 2021; Liu et al., 2022b; Fu et al., 2022) have explored shape editing guided by semantic tags or language descriptions, 3D generation is still challenging due to the lack of 3D data and suitable architectures. Based on the diffusion model, DreamFusion (Poole et al., 2022) proposes to solve these problems with a pre-trained text-to-2D model. Another branch of works reconstruct the 3D objects from single-view images (Bednarik et al., 2018; Tsoli et al., 2019; Golyanik et al., 2018; Wang et al., 2018b; Yuan et al., 2021c; Li and Kuang, 2021) or multi-view images (Huang et al., 2018; Choy et al., 2016; Xie et al., 2019; Wang et al., 2021a), termed Image-to-3D. A new branch of multi-view 3D reconstruction is Neural Radiance Fields (NeRF) (Mildenhall et al., 2021) for implicit representation of 3D information. The 3D-3D task includes completion from partial 3D data (Wang et al., 2022) and transformation (Bai et al., 2015), with 3D object retrieval as a representative transformation task.

3. Speech

Speech synthesis is an important research area in speech processing that aims to make machines generate natural and understandable speech from text. Methods of traditional speech synthesis include articulatory (Kröger, 1992; Shadle and Damper, 2001), formant (Allen et al., 1979; Seeviour et al., 1976), concatenative synthesis (Olive, 1977; Moulines and Charpentier, 1990), and statistical parametric speech synthesis (SPSS) (Kawahara, 2006; Morise et al., 2016). These methods have been widely studied and applied, e.g., formant synthesis is still used in the open-source NVDA (one of the leading free screen readers for Windows). However, these generated speeches are identifiable from the human voice, and artifacts in synthesis speech reduce intelligibility.

Early works (Ze et al., 2013; Qian et al., 2014; Fan et al., 2014; Zen, 2015; Zen and Sak, 2015) consist of three modules: text analysis, an acoustic model, and a vocoder. WaveNet (van den Oord et al., 2016) is a revolution within speech synthesis which can generate the raw waveform from the linguistic features. To improve the quality of speech and diversity of voices, generative models are introduced in speech synthesis, such as GAN (Goodfellow et al., 2014). Compared with GAN, diffusion models do not require a discriminator, making training more stable and simple. Therefore, the works of speech synthesis adopt diffusion models, becoming a rising trend. A branch of works (Huang et al., 2022b; Xiao et al., 2021; Lam et al., 2022; Chen et al., 2022) focuses on efficient speech synthesis, in which different ways are adopted to reduce the generated time by accelerating inference, such as combining the schedule and score networks for training, jointly trained GAN. Another branch of study (Chen et al., 2021; Shi and Wu, 2022; Mittal et al., 2021; Rouard and Hadjeres, 2021) concentrates on end-to-end models, which directly generate waveform from text without any intermediate representations. A fully end-to-end model not only simplifies the training and inference, but also reduces the demand for human annotations. The branch of diffusion-based speech synthesis is not limited to the two mentioned above, such as speech enhancement and guided speech synthesis.

4. Graph

Graphs are ubiquitous in the world, which aid in visualizing and defining the relationships between objects in a wide range of domains, from social networks to chemical compounds. Graph generation, which creates new graphs from a trained distribution that is similar to the existing graphs, has received a lot of attention.

Traditional graph generation works (Watts and Strogatz, 1998; Albert and Barabási, 2002; Leskovec et al., 2010) create new graphs with specific features that are related to the hand-crafted statistical features of real graphs , which simplifies the process but fails to capture relational structure in complex scenarios. With the successes of deep learning algorithms, researchers have begun to apply them to graph generation, which, unlike the traditional methods, can be directly trained by real data and automatically extract features. Among them, works (You et al., 2018; Liao et al., 2019; Dai et al., 2020) based on autoregressive model create graph structures sequentially in a step-wise fashion, which allows for greater scalability but fails to model the permutation invariance and is computationally expensive. Simultaneously, One-shot models (Liu et al., 2018; Madhawa et al., 2019, 2019) such as VAE and flow are incapable of accurately modeling structure information because of ensuring tractable likelihood computation. Although graph generation (De Cao and Kipf, 2018; Jin et al., 2018; Maziarka et al., 2020) based on GAN sides step likelihood-based optimization by using a discriminator, the training is unstable.

Recently, there has been a surging interest in developing diffusion models for graph-structured data. EDP-GNN (Niu et al., 2020) is the pioneering to show the capability of the diffusion model in the Graph generation, with the goal of addressing non-invariant properties. After that, On the one hand, diffusion-based works (Jo et al., 2022; Vignac et al., 2022; Haefeli et al., 2022; Luo et al., 2022; Huang et al., 2022a) focus on realistic graph generation, which produces graphs that are similar to a given set of graphs. On the other hand, (Shi et al., 2021; Xie et al., 2021; Anand and Achim, 2022; Zaman et al., 2022) concentrate on goal-directed graph generation, which generates graphs that optimizes given objects, like molecular and material generation.

5. Others

There are also other interesting tasks generating content in different modalities, e.g., music generation (Ji et al., 2020) and lip-reading (Fenghour et al., 2021). A typical music generation system can be categorized into three representation levels (from top to bottom), which generates score, performance, and audio, respectively (Ji et al., 2020). With the development of deep learning, music generation introduces various methods for higher music quality, e.g., MusicVAE (Roberts et al., 2018), MuseGAN (Dong et al., 2018) and transformer in (Huang and Yang, 2020). Music generation inspires and accelerates the development of computer-assistant composition software, including Magenta project from Google and Flow Machine project from Sony Computer Science Laboratories. A Lip reading task transforms visual inputs of lip movement to decoded speech (Fenghour et al., 2021), and has also shown impressive advances thanks to improved corpora and architectures.

Industry Applications

Undoubtedly, AIGC has gone viral on social media since 2022. For example, users are active in sharing their experience of using ChatGPT for having an interactive conversation or Stable diffusion for generating images with a text prompt. However, this hype is expected to dwindle if AIGC cannot be used for practical applications in the industry to demonstrate its value. Therefore, we discuss how AIGC might influence various industries.

AIGC is changing the paradigm of education by assisting in teaching and learning. Generative AI carries transformative potential in teaching, with the application ranging from course materials generation to assessment and evaluation (Zentner, 2022; Pettinato Oltz, 2023). Simultaneously, applications of generative models have begun to influence how students learn (Tangermann, 2023; Baidoo-Anu and Owusu Ansah, 2023).

Generative AI technologies can provide educators with creation of personalized tutoring (Zentner, 2022), designing course materials (Pettinato Oltz, 2023), and assessment and evaluation (Zentner, 2022; Baidoo-Anu and Owusu Ansah, 2023). A unique foreign language teaching product for young children using generative technologies such as ChatGPT can attract children’s attention, motivate them, and provide a fun learning environment. Higher education needs to embrace the use of AI in higher education, which can create more engaging, effective, and efficient learning experiences for students (Zentner, 2022). One of the primary benefits of generated AI course material generation is that it can save teachers time and effort by automating the process of creating and updating course material. In addition, ChatGPT could significantly reduce the workload of law school instructors, freeing up time to increase academic productivity or develop more complex teaching skills (Pettinato Oltz, 2023). The benefits of ChatGPT in promoting teaching include but are not limited to facilitating personalized and interactive learning. However, some limitations of ChatGPT, such as generating incorrect information, exacerbating existing biases in data training, and privacy issues, can also appear (Baidoo-Anu and Owusu Ansah, 2023). Overall, addressing these challenges requires collaborative efforts from policymakers and educators to provide recommendations or guidance for the appropriate use of generative AI tools.

Moreover, generative AI technologies can help students write essays (Tangermann, 2023), at-home tests or quizzes (Tangermann, 2023), comprehend certain theories and concepts, and different language essays and papers in academic issues (Zentner, 2022; Baidoo-Anu and Owusu Ansah, 2023). Chatbots can provide students with 24/7 support, allowing them to get the help they need when they need it. With the ability to correct grammar, suggest improvements, and identify weak areas, chatbots like ChatGPT can provide students with immediate feedback on their writing, helping them to learn from their mistakes and improve their writing skills over time. This not only saves students time but also helps them to become better writers (Yadav, 2023). According to a survey conducted by an online course provider, 89% of students use ChatGPT to complete their homework, with 50% using it for essays and 48% using it for at-home tests or quizzes (Tangermann, 2023). Additionally, generated AI can tailor the course material to individual students’ needs, such as learning style and pace, which has the potential to improve student engagement and learning outcomes. ChatGPT can also help students comprehend certain theories, concepts, and different language articles, making them work more effectively (Zentner, 2022; Baidoo-Anu and Owusu Ansah, 2023). There are also challenges and concerns associated with generated AI course material generation, including the generated material’s quality, and the possibility of bias in the data used to train the AI. As a result, before using generated course material in an educational context, it is critical to evaluate and validate it carefully (Dawn Gilmore, 2023).

With the use cases mentioned above, AIGC has the potential to revolutionize education by improving the quality and accessibility of educational content, increasing student engagement and retention, and providing personalized support for learners. With the continuous advancements in AI, AIGC is poised to become an integral part of the education industry, offering students a more engaging, accessible, and personalized learning experience.

2. Game and metaverse

Most users may not resonate with one-size-fits-all content in the game and metaverse, where personalization yields the best experience. Although games and metaverse provide users with virtual worlds, the content represents the character and personality of users. Generative AI makes that possible, which not only allows users to customize their avatars but also provides diverse scenarios and storylines, making the experience more immersive (Ratican et al., 2023; Chen, 2022; Plut and Pasquier, 2020; Ratican et al., 2023).

AI Dungeon powered by GPT-3 model allows users to generate an open-end story navigated by text, where generative AI will produce new events as the response to the different actions of users, creating a one-of-a-kind and unexpected gameplay experience (Team, 2020). Horizon Worlds, one of the most popular games, allows you to wander into the virtual world related to content consumption and creation. In Horizon, users will have more control over how they want to tailor their online experience to meet their individual needs. Specifically, users can design their unqiue avatars and scenes using gizmos that include pre-built object and avatar properties (Meta, 2022). Moreover, the visual novel game Traveler concentrates on generating gorgeous scenarios to present users with visual impact, in which you will embark on a journey through a diverse world. When a player explores the game Traveler, the player will be exposed to magnificent visuals and immersive soundscapes. As each scene is unique, the content can range from dark forests to bustling cities, all crafted by generative AI (Games, 2023).

Although the term “metaverse” has become a buzzword recently, in the real world, the virtual space created by the game may serve as the portal to the metaverse (Ratican et al., 2023; Ning et al., 2021). Roblox, a sandbox game platform, first included the concept of “Metaverse” in its prospectus and made its market value soar, where players can create their own world beyond their imagination (Ning et al., 2021). Virtual concert singers have a more comprehensive range of musical styles and talents. Audiences can choose their own favorite styles and even idols, providing them with a more diverse and personalized concert experience. Travis Scott, a well-known American rapper and producer, performed a historic concert inside the Fortnite game, and his avatar guided the players to experience different scenes, ranging from underwater to outer space (Webster, 2020). The University of California, Berkeley, presented its commencement ceremony in ‘Minecraft’, a popular computer game. In the Minecraft game, students and alumni built a copy of campus using generative AI technology, allowing thousands of graduating seniors using their avatars from around the world to attend the event (Rosato, 2020). Overall, AI has played a significant role in the evolution of the game and metaverse, and its use continues to grow as technology improves and becomes more accessible.

3. Media

With the ubiquitous growth of generative AI technologies, they play a rising role in media and advertising. AIGC not only promotes the diversity of media, which provides a better experience for audiences, but it also enables media practitioners more efficient in their work (Miroshnichenko, 2018; Latar, 2018; Fırat, 2019; Zhang, 2022; Vakratsas and Wang, 2020; Campbell et al., 2022; Deng et al., 2019).

The media powered with AIGC enables more diversified content and ways of reporting, changing the media mode of production and organizational structure (Zhang, 2022; Kim and Kim, 2017; Marconi, 2020). AIGC can be applied to a variety of applications in media, such as writing robots, news anchors, and caption generation. Traditionally, media outlets have relied on expert journalists to write new articles and reports, which requires a significant amount of energy and time, resulting in a limited number of articles. Moreover, the timeliness of the news is critical, and the news may be eclipsed after an hour. Generative AI can greatly assist journalism by using text generation technologies to make journalism more efficient and responsive (Miroshnichenko, 2018). Associated Press applies these technologies to generate roughly 40000 stories a year and its articles on company earnings increase from 1200 to 14800 (Gruber, 2022). The Quakebot, a robot reporter of Los Angeles Times News, only takes three minutes to complete a related article after the Los Angeles earthquake (Oremus, 2014). Bloomberg News, an international financial media company, launched Buttetin in 2018, with the goal of providing personally storied whose one-sentence summaries are generated by chatbots (Willens, 2018).

AI news anchors have emerged as a result of the deep integration of generative AI in the media (Wang, 2021; Xue et al., 2022; wang2021research). AI news anchors, combined with real anchors, make the way to spread information more diverse. AI news anchors can broadcast news based on the text, whose appearance and expression imitate the real anchor. China’s state news agency Xinhua and Chinese search engine, Sogou, have developed AI news anchors with different profiles and languages. The most impressive is the 3D AI news anchor Xin Xiaowei, whose broadcast form can be presented in all directions from various angles, significantly improving the sense of three-dimensionality and layering (McFarland, 2022). Additionally, Korea’s cable channel MBN has created the AI news anchor AI Kim, who can quickly respond to various emergencies and even report all day (So-Yeon, 2020). To help the hearing-impaired people to get more information about international sporting events and Beijing Winter Olympics, an AI-driven sign language service provider, namely Ling Yu, was developed by giant tech Tencent, where AIGC tasks are used including 3D digital human modeling, machine translation, image generation, and speech-to-text (Yuanjie, 2022). Moreover, Migu, a Chinese business dedicated to providing digital information, can offer smart subtitle functions to the live broadcast of Beijing Winter Olympics. Therefore, people with hearing impairments can watch the live broadcast of the sports event, which makes them more immersive.

4. Advertising

Various AI applications have transformed the advertising industry, giving advertisers powerful tools to create innovative and engaging content that connects to consumers at a deeper level (Vakratsas and Wang, 2020; Deng et al., 2019; Guo et al., 2021). Among the various applications of AI in advertising, AIGC is particularly influential by allowing advertisers to create personalized and attractive content that resonates with individual consumers. A creative advertising system (CAS) aligns with the principles of AI for generating and testing advertising ideas, which helps aspiring and mature creators understand that creativity is not an elite privilege, but rather a systematic process that can be assisted through data and computation (Vakratsas and Wang, 2020). Implementing programmatic advertising has not fully utilized self-generating technology, resulting in different consumers being exposed to the same content. Fortunately, a personalized advertising copy intelligent generation system (SGS-PAC) can automatically personalize advertising content to meet individual consumer needs (Deng et al., 2019). Advertising posters are a common form of information display used to promote products. Another intelligent system, Vinci, supports the automatic generation of advertising posters (Guo et al., 2021). By inputting product images and slogans specified by users, Vinci uses deep-generation models to generate beautiful posters.

In addition, Brandmark.io is an AIGC-based tool that automatically generates logos for businesses. The tool creates multiple logo variations based on the user’s preferences and specifications. Advertisers can purchase and use the logos created by the tool for their businesses, making it an easy and cost-effective solution for logo design (kapadia, 2023). By GAN that forces output to include specific keywords, the approach automates product listing generation likely to attract potential buyers (Martinez and Kamalu, 2018). It enhances users’ marketing efforts on peer-to-peer marketplaces. Moreover, technological innovations have provided digital and automated tools to the advertising industry, but have also allowed advertisers to automate the production of “synthetic advertising”. As reported by (Shah et al., 2020), AIGC has transformed the advertising industry by enabling advertisers to create highly personalized and engaging content at scale while saving time and resources. We expect to see even more innovative and influential applications of generated AI in advertising.

5. Movie

It is interesting to see how technology now affects almost every step of movie creation. Research has led to the development of computer-based surroundings that help with editing, labeling, video retrieval, and many more (Davis, 1994; Gordon and Domeshek, 1995; Parkes, 1988, 1989b, 1989a; Sack and Davis, 1994; Swanberg et al., 1993; Tonomura et al., 1994; Ueda et al., 1993; Nack, 1996). To start with, AI-powered screenwriting software has significantly impacted the movie-making process. AI has created a new movie experience by integrating visual effects (VFX) (Momot, 2022), improved sound effects (SFX), and new viewing platforms. The 4K, IMAX, and 3D movies, as well as animations, are highly impacted by them.

The script forms the foundation of how a movie will fare at the ticket counters. AI-generation devices store and compute massive amounts of data to create “ideal” scripts (Anantrasirichai and Bull, 2022). AI software is also employed to rework old screenplays into polished versions that are then analyzed and improved by the director and writer. It goes beyond just developing and analyzing current scripts. Jasper AI (Ruby, 2023) and Scalenut (Pri, 2022) are two examples of AI scriptwriters.

Movies are given visual effects (VFX) to increase spectator appeal. They combine original images with real video to create engrossing, realistic, contextual depictions that may include digital surroundings, de-aging, and many more. The VFX team of The Curious Case of Benjamin Button (Sydell, 2009) put up two arrays of cameras in a bright room and utilized the MOVA Contour reality capture technology (DeMott, 2006) to construct a three-dimensional database of the hero’s facial expressions. The VFX team next developed high-resolution 3D models of the lead character at various ages and lastly employed AI to manipulate the data retrieved from the three-dimensional database to cause the head models to age. The outcome was convincing and garnered widespread acclaim across the movie community. In movies like Blade Runner 2049 (Marshall, 2018) and Gemini Man (ROTTENBERG, 2019), the VFX team tracked and recorded the protagonist’s facial data with the help of motion capture technology (Adobe, 2022). They then rebuilt the 3D facial data in a computer and further polished it to accomplish age deduction. Although this approach takes expensive technology and a vast amount of money, it is incredibly precise and adaptable.

AI is expanding the limits of amusement by bringing back actors who have passed away in movies. Deepfake technology uses computer visuals and AI to produce incredibly amazing and lifelike fake videos of actual people or made-up characters. Fast and Furious 7 (WIKIPEDIA, 2013) used VFX and vintage video to bring back Paul Walker after he passed away in 2013, using the actor’s visage transferred onto his sibling.

In addition to the visual effects, subtitles also play a vital role in viewers’ experience. For the benefit of viewers having hearing impairments, automated subtitles for the deaf and hard of hearing, also known as SDH (ai media, 2017), include textual transcriptions of speech, speaker changes, and background noise. These benefits substantially increase how much money movies make. Natural Language Processing (NLP), a kind of AI that focuses on deciphering spoken language, offers multilingual subtitles in movies. To produce automated subtitles, Rev (@rev, 2023), a cloud-based program, is widely used by movie lovers. AI has also revolutionized the task of speech prediction in silent movies. AI-generated speech synthesis can narrate silent movies and dub movies into multiple languages. Deep learning systems trained on massive human audio samples can produce natural-sounding voiceovers. Features like LPC (Linear Predictive Coding) (Ephrat and Peleg, 2017) and mel spectrograms (Ephrat et al., 2017) generate high-quality intelligible audio through conversion. Recently, a Tacotron2 model (Shen et al., 2018) variation for the video-to-audio synthesis was proposed by (Prajwal et al., 2020a). (Yadav et al., 2021) suggests an efficient stochastic model that produces endless high-quality audio patterns for a specific silent video, thus effectively encapsulating the multimodality of the speech prediction issue. Along with the visual and sound effects, we have Colourlab.Ai (@ColourlabAI, 2023) for color grading, Descript (@descript, 2022) for video editing, and many more tools continuously making waves in the movie industry.

6. Music

AIGC also makes it to the music industry with notable developments (Yang and Nazir, 2022). AI can not only spot patterns and trends in vast data sets that are challenging for humans to notice, but also allows amateur musicians a cutting-edge technique to enhance their creative process, which is a fantastic opportunity. The fusion of AI technology into music is a new trend that many experts, researchers, musicians, and record companies are exploring (McFarland, 2023). Many utilize AIGC to create entirely new music, while some software edit compositions in the style of various composers.

Music industries are anticipating significant expenditures in this field, whether it is because of using AI to compose music or to help musicians. A fantastic illustration of an AI melody generator is Google’s Magenta project (Hutson, 2017), (Alaeddine and Tannoury, 2021). IBM’s Watson Beat is one more example. For composing an original song, it makes use of AI and machine learning (Chaney, 2018), (Frid et al., 2020). 2016 saw the successful creation of text-to-speech (TTS) recordings and recordings that resembled music by DeepMind researchers (Team, 2017), (Yang et al., 2017). AI is also vastly used for the processing and improvement of digital audio. LANDR, an incredible AI-powered creative tool that enables musicians to get their music on several streaming services like Spotify and Apple Music, is one such service. A significant problem known as “writer’s block (Sexton, 2023)” frequently confronts lyricists. But thanks to AI, it’s no longer a problem now. Nowadays, many musicians employ AI to create new lyrics for their songs (Maya Ackerman, 2022). GPT-2 (AI, 2019), (Zhou et al., 2023), a text-generating tool, has been created by OpenAI, an AI technology firm. Not only can this remarkable text generator produce authentic news, but it can also write lyrics for Beatles songs and music from all other genres. However, AI is not just capable of producing text; it can also create original soundtracks and melodies. The Sony CSL flow machine offers assistance to artists so they can develop original music based on their ideas (Rogerson, 2020). One of the most well-known AI tools for writing unique music is called AIVA (Juillet, 2021). To produce a unique track, the user first chooses a pre-set style and then modifies a variety of variables, such as the key, instrumentation, time signature, etc. AIVA can deconstruct all the intricate auditory information saved into discrete characteristics while reading hundreds of musical compositions by renowned musicians like Bach and Mozart (Drott, 2021). These qualities may then be interpreted again and used to produce a completely new musical work. Apart from the tools mentioned above, there are so many other applications that made a significant impact on the music industry, such as the iOS-based tool Amadeus Code (Blog, 2019), the cloud-based platform Amper (Heater, 2022), Ecrett Music (Berry, 2019), etc.

7. Painting

From offering automatic painting tools to encouraging creative experimentation, AIGC is revolutionizing the painting industry in many ways. AI programs can analyze pictures to produce color schemes, patterns, and textures that can make artwork. The automatic drawing tools generated using these algorithms are able to apply these patterns and textures to produce distinctive and intricate works of art (Zou et al., 2021). AI can also analyze a person’s preferences, interests, and style to create customized artwork. Empowering artists to create art specifically suited to their preferences and interests can increase their appeal and value.

The artwork created by MidJourney under the title ‘Space Opera Theatre’ earned first place in the Colorado State Fair Art Competition (DelSignore, 2022), demonstrating the capability of AI painting tools to produce excellent pieces of art. Midjourney is an excellent AI image generator with comprehensive functions, which is used by many artists to generate inspiration. The creation of various art forms by generative AI, such as abstract painting generation (Li et al., 2020b), Chinese shanshui painting (Zhou et al., 2019b), and Chinese ink paintings (Chung and Huang, 2022), undoubtedly promotes the advancement of painting. Moreover, AIGC can assist in conservation and restoration (Yu et al., 2022a), (Hu, 2022). AI algorithms are capable of analyzing and repairing ruined artwork. These algorithms make it simpler for conservators to return the artwork to its initial state by detecting and removing dust, scratches, and other flaws.

For non-professionals unfamiliar with drawing or animation, AIGC is also very helpful because it enables them to produce high-quality visual effects. By adding additional constraints to the diffusion model, ControlNet (Raieli, 2023) can increase the variability of the produced images. It can describe the generated images along with those other constraints of border drawing, depth information, Hough line map, normal map, and posture estimation. AIGC has also started a new era of collaborative artwork (Chang et al., 2022; Bublitz et al., 2019). AI algorithms can create collaborative paintings that involve multiple artists working together. These algorithms can analyze the styles of each artist and produce a unified style that incorporates elements from all of the artists’ works.

8. Code development

Generative AI can contribute to the field of code development (Sun et al., 2022; Guo et al., 2022; Guo, 2021; Jahić et al., 2019), where AIGC can create code without the need for manual coding. The work by (Sun et al., 2022) explores the interpretability requirements of generative AI for code and demonstrates how human-centered approaches can drive the development of explainable AI (XAI) technologies in new domains. To improve testing efficiency and increase test coverage, it is particularly important to generate high-quality test cases automatically (Guo et al., 2022; Guo, 2021). In order to optimize the efficiency of data engineering, a novel software engineering approach based on neural networks for dataset augmentation can be designed (Jahić et al., 2019). One of the popular applications in AI-generated code is Github’s Copilot, an AI tool jointly developed by GitHub and OpenAI. Users can automatically complete code through GitHub Copilot using software development tools (Dohmke, 2022). Moreover, AI-generated technology can also assist in code refactoring, which improves existing code without changing its original functionality. This can shorten the time for developers to refactor and improve the quality of the code. A popular code refactoring tool is DeepCode (Ingle, 2023), an AI-supported code review tool that can inspect your code and provide suggestions for improvement. In addition, AIGC can also make an impact on the e-commerce and finance industries (Houde et al., 2020). E-commerce platforms such as Amazon, JD.com, and so on can use AI-powered customer service to provide shopping guide services to customers, thereby saving costs for enterprises. Financial companies can use virtual investment advisors to advise customers on securities account opening, financial investment, and other related services.

9. Phone apps and features

Numerous AIGC applications have emerged as fun-oriented mobile apps, typically in the form of image and video editing. Photoshop is traditionally a common tool for image editing, but manual work is time-consuming and can result in unnatural or unrealistic output. In addition, video editing involves analyzing each video clip and making editorial decisions based on both the audio and visual content. This process is time-consuming because the video is a time-based, dual-track medium that requires careful consideration of every frame. Fortunately, some work (Bar-Tal et al., 2022; Soe, 2021; Argaw et al., 2022) has explored the utilization of AI technologies behind AIGC, to the image or video editing, making the applications in AIGC such as face swapping and digital avatar possible.

Some popular applications based on face swapping are gaining widespread popularity on the Internet. This technology uses advanced AI technologies to analyze and swap people’s faces with their favorite celebrities or anyone else in seconds, making it easier and faster to use compared to traditional PS technologies. VanceAI, Voila AI Artist and FaceAPP are leading figures, with FaceApp being recognized as the best facial photo editing App, winning numerous awards, and being downloaded by over 500 million users and counting (Salia, 2021). Another popular application is voice-changing technology. This technology can adjust the pitch, timbre, speech rate, and other characteristics of the human voice to change the quality of the human voice. MagicMic (William, 2021) and Voicemod (Kenyon, 2022) are two popular applications for real-time voice modification and soundboard operations, which people can use to change their voices for creating fun content, live streaming, or other purposes, enhancing the enjoyment of communication between people.

In addition, another technological trend is to transform individuals into virtual characters, thereby increasing entertainment value. virtual characters are digital avatars of people in a virtual world, they can be partial replicas of real people or even completely digital versions. Apple’s first “digital avatar” technology, Animoji, focuses mainly on generating preset cartoon and animal characters and does not support custom generation (Tillman, 2022). The second generation of “digital avatar” technology represented by iPhone’s Memoji and Xiaomi’s Mimoji started to support personalized avatar customization, which offers a variety of options, starting from hairstyle, eyes, nose, dresses, etc (Statt, 2019). This upgrade allows users to create an avatar that not only can track their facial movements, but also look like them. Besides that, the created avatars can also be posted as comments in WeChat or Facebook chats, giving users a more personal way to express themselves on social media. Since then, digital avatar technology has become one of the standard features of smartphones among various smartphone manufacturers.

10. Other fields

Beyond the above fields, AIGC is expected to have applications in more fields. For example, the design and development of a novel drug are complex, costly, and time-consuming. On average, it takes around $3 billion and more than 10 years for a new drug to be accepted by the market (Kelder and Parker, 2021). This motivates using AIGC to accelerate the drug discovery process and reduce costs. In 2018 DeepMind created AlphaFold (Ruff and Pappu, 2021), which can accurately predict the structure of proteins and has been considered a milestone for drug discovery and fundamental biology research. Its updated version AlphaFold2 was released in 2020 and had higher accuracy than the former. ProteinMPNN(Dauparas et al., 2022), designed by Justas Dauparas, can design protein sequences for specific tasks, generating entirely new proteins quickly in just a few seconds. Besides directly exploiting the generated content, AIGC can also help workers in various fields improve their efficiency. For example, in medical consultation, the patient can rely on chatbots for basic medical advice, while turning to the doctor only for more severe cases. In manufacturing design, it is possible to combine AIGC with the widely used computer-aided design system to minimize the repetitive effort so that the designer can focus on the more meaningful part.

Challenges and outlook

Even though AIGC has shown remarkable success in generating realistic and diverse outputs across various domains, there are still numerous challenges in real-world applications. Except for requiring a large amount of training data and compute resources, we list some of the most significant challenges as follows.

Lack of interpretability. While AIGC models can yield impressive outputs, it remains challenging to understand how the model arrives at the outputs. This is especially a concern when the model generates an undesirable output. Such a lack of interpretability makes it difficult to control the output.

Ethical and legal concerns. The AIGC model is prone to data bias. For example, a language model mainly trained on the English text can be biased toward western culture. Copyright infringement and privacy violations are the underlying legal concerns that cannot be ignored. Moreover, the AIGC model also has the potential for malicious use. For example, students can exploit these tools to cheat on their essay assignments, for which AI content detectors are desired. AIGC models can also be used for distributing misleading content for political campaigns.

Domain-specific technical challenges. At the current and in the near future, different domains require their unique AIGC models. Each domain is still faced with its unique challenges. For example, Stable Diffusion, a popular text-to-image AIGC tool, occasionally generates output that is far from what the user desires, such as drawing humans as animals, one person as two people, etc. ChatBot, on the other hand, makes factual mistakes occasionally.

2. Outlook

Despite its unprecedented popularity, generative AI is still in its early stage. Here, we present how AIGC might evolve in the near future.

More flexible control. A major trend of AIGC tasks is to realize more flexible control. Taking image generation as an example, early GAN-based models can generate images of high quality, but with little control. The recent diffusion models trained on large text-image data enable control through text instruction. This facilitates the generation of images that better match the users’ needs. Nonetheless, current text-to-image models still require more fine-grained control so that the images can be generated in a more flexible manner

From pertaining to finetuning. Currently, the development of AIGC models like ChatGPT focuses on the pretraining stage. The corresponding technology is relatively mature; however, how to fine-tune these foundation models for the downstream tasks is an under-explored field. Different from training a model from scratch, the goal of finetuning needs to trade-off between the foundation model’s original general capability and its adaptation performance on the new task.

From big tech companies to startups. At present, AIGC technology has mainly developed big tech companies, like Google and Meta. With the support of big tech companies, some startup companies have emerged to show high potentials, like OpenAI (supported by Microsoft) and DeepMind (supported by Google). With the focus transition from core technology development to applications, more startup companies are expected to emerge due to increasing demand.

Discussion: investment, bubble and job opportunities. Technology-wise, there is no doubt that AIGC has made significant progress in the past few years. When a transformative technology emerges, the market tends to be over-optimistic about its potential applications and future growth, which also applies to generative AI. According to PitchBook (see Fig. 21), the funding for generative AI from venture capital (VC) increased significantly in the last two years. Some critics have concerns that generative AI might be the next bubble. One of their main concerns is that most AIGC tools are mainly playful instead of practical. For example, text-to-image models are fun to play with, but how they might generate revenues remains unclear. It is difficult to predict how generative AI might evolve. However, the authors of this work believe that generative AI is unlikely to become the next bubble considering it is a relatively new and rapidly growing field with many potential applications. There is also a hot debate about whether generative AI will replace humans, causing the loss of numerous job opportunities. On the other hand, generative AI can also create new job opportunities for individuals with skills on AI research and implementation skills. The industries that benefit from the power of AIGC might also boom and generate more job opportunities.

References