CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models
Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Libiao Peng, Jiaming Yang, Xiyao Xiao, Sahand Sabour, Xiaohan Zhang, Wenjing Hou, Yijia Zhang, Yuxiao Dong, Jie Tang, Minlie Huang
Introduction
Large language models (LLMs) (Touvron et al., 2023a, b) have transformed the shape of not only research but also industrial applications (Brown et al., 2020; Ouyang et al., 2022; OpenAI, 2023). Designed as universal task assistants, these models have demonstrated unprecedented capabilities in intent understanding, instruction following, and task solving in a wide array of applications (Google, 2023; Anthropic, 2023). They have shown great efficiency and effectiveness in solving mathematical problems (Zhou et al., 2023a), reasoning (Wei et al., 2022), code debugging (Zheng et al., 2023), and even image understanding and generation (Betker et al., 2023).
Despite their great success, existing LLMs are still incompetent in accomplishing social goals, for instance, establishing long-term social connections with humans or providing effective emotional support for humans (Liu et al., 2021; Zhou et al., 2023b). Since LLMs are typically designed to solve various tasks, the dialogue data used in training often consists of exchanges in short turns, unlike daily life conversations. However, this line of research can date back to the early stages of artificial intelligence. In 1966, MIT built a chatbot ELIZA (Weizenbaum, 1966) for psychotherapy, which simulated a psychological counselor by generating conversations via linguistic templates and keyword detection. Prior to the emergence of ChatGPT, open-domain dialogue systems, which aim to build social connections with humans through conversational interactions, have attracted tremendous research efforts and have been advanced significantly. Some examples of such systems include BlenderBot (Roller et al., 2021; Shuster et al., 2022), Meena (Adiwardana et al., 2020), LaMDA (Thoppilan et al., 2022), EVA (Zhou et al., 2021; Gu et al., 2023), and Plato (Bao et al., 2021, 2022). These models can deliver very natural, human-like conversations, some of which are even regarded as having their own persona and consciousnesshttps://www.youtube.com/watch?v=NAihcvDGaP8.
As illustrated in Maslow’s hierarchy of needs(Maslow, 1974), the feeling of love and belonging serves as a vital need for humans, which could be fulfilled by conversational engagement with desired social entities. Recognizing the significance of social connections, the development of LLMs tailored for social interaction emerges as a crucial endeavor. To this end, we meticulously define the concept of “character” to represent a social entity and introduce a new task named Character-based Dialogue (CharacterDial). As depicted in Figure 1, CharacterDial allows users to specify and create profiles for their preferred characters, ranging from fictional figures like “Sun Wukong” to realistic personalities such as “Lu Xun”. The conversational AI system then adopts the worldviews, values, and common facts of these profiles, presenting a tailored character that facilitates engaging and extensive conversations with users.
In pursuit of character-based dialogue, a straightforward approach involves role-playing with LLMs in a tuning-free manner (Yu et al., 2022) where the LLMs are instructed to follow specified character descriptions. However, this approach encounters significant challenges in accurately reflecting the intrinsic relationship between the character profile and the dialogue content. As illustrated in Figure 2, when applied to multi-turn interactions in CharacterDial, LLMs show inferior performance in consistency, human-likeness, and engagement, especially in long-range dialogue turns. Another relevant field of study is persona-based dialogue (Zhang et al., 2018), which aims to generate responses based on superficial interlocutors’ personal information such as genders, names, and hobbies. However, this approach has limitations in fully capturing the complex attributes (e.g., identities, interests, etc.) and behaviors (e.g., linguistic features, emotional expressions, etc.), which are essential in constructing a three-dimensional character (Wardhaugh and Fuller, 2021). Consequently, it often fails to generate dialogues that reflect a character with its own unique style and vivid personality. As depicted in Figure 3, character-based dialogue can be viewed as an extension of persona-based dialogue, presenting a more comprehensive task setting. Recently, Character.AIhttps://character.ai/ has emerged as a frontrunner in the commercial implementation of character-based dialogue. Nevertheless, its foundational model is proprietary and not open to the academic and industrial communities. This limitation highlights a gap in accessible resources for extensive research and development in character-based dialogue systems.
In this paper, we propose CharacterGLM, a series of large language models for customizing AI characters to deliver consistent, human-like, and engaging conversations. Accordingly, we introduce a
novel task of generating character-based dialogue (CharacterDial), which enables us to customize virtual conversational AI characters. We crowdsourced a large-scale Chinese CharacterDial corpus from various sources, covering a diverse range of character categories and dialogue topics. Each dialogue session establishes a specific character, complemented by a profile containing their attributes and behaviors. We will release a portion of this corpus to public use, which consists of 1,034 high-quality dialogue sessions spanning 250 charactershttps://github.com/thu-coai/CharacterGLM-6B. We then develop CharacterGLM, a collection of large language models building upon ChatGLM (Du et al., 2022; Zeng et al., 2022) with carefully designed training and self-refinement methods, to support flexibly customizing characters for CharacterDial. CharacterGLM models vary in size from 6B to 66B parameters. We release the 6B version to the research communityhttps://huggingface.co/LingxinAI/CharacterGLM-6b, while access to other versions is provided through our APIhttps://maas.aminer.cn/dev/api#characterglm.
Design Principle of CharacterGLM
The development of conversational AI characters centers on creating a virtual conversational partner that is realistic, trustworthy, and engaging. This demands a thorough comprehension and imitation of human communications, particularly in the form of text-based interactions. In this context, we analyze the human traits that influence conversational expressions and divide them into two primary categories: attributes and behaviors. Attributes are predominantly reflected in the content of responses, whereas behaviors focus on more tone and style. Additionally, we assess the effectiveness of character-based dialogue from three aspects: the degree to which conversational expressions adhere to human traits (consistency), the naturalness of the conversational style in reflecting human-human interactions (human-likeness), and the extent to which the dialogue can attract and engage users (engagement).
Humans are multifaceted beings with various attributes that represent static or gradually evolving features. These attributes provide essential background information for replicating an individual as a conversational AI character, significantly influencing the character’s reactions and interactions (Grice, 1975; Justine Cassell and Churchill, 2000). For instance, a character’s identity can be determined by its cultural background or occupation, and its interests may include hobbies and preferences. Viewpoints, on the other hand, can guide its morals, values, and beliefs. By integrating these attributes, conversational AI characters can more accurately mimic the way that humans draw on their unique information to manage communication. In CharacterGLM, we consider seven primary categories of attributes: a) Identities: encompassing name, gender, age, birth date, occupation, residence, family composition, belongings, etc. b) Interests: including preferred and disliked items. c) Viewpoints: covering worldviews, life philosophies, and values. d) Experiences: containing past and present experiences. e) Achievements: such as awards and honors. f) Social relationships: detailing connections with parents, teachers, classmates, etc. g) Other: comprising skills, specialties, etc.
Behaviors in conversational AI characters are represented by dynamic elements such as linguistic features, emotional expressions, and interaction patterns, which are crucial in shaping realistic dialogue contexts(Pickering and Garrod, 2004). For instance, a character defined as “elderly” might use more formal language, while a “teenager” might employ current slang. Language production in humans is not only a matter of conveying information but also a form of action that is affected by one’s social and psychological state (Austin, 1975). Incorporating these aspects into the behavior of conversational AI characters allows for a more natural and human-like dialogue flow, which is critical in maintaining the user’s curiosity when interacting with AI characters. In CharacterGLM, we consider linguistic features, including a person’s catchphrase, dialect, stylistic features, frequently used words and sentences, etc. In addition, we also consider personality as an important factor in shaping response, such as gentleness and coldness.
2 Features of AI Characters: Consistency, Human-likeness, and Engagement
Character consistency refers to the need for the conversational AI character to display a stable set of attributes and behaviors during interactions. Consistency is essential for believability and trust in human conversations and, by extension, in conversational AI interactions (Nass et al., 1994). According to the psychological concept of personality consistency (John et al., 1999), individuals tend to exhibit stable behavior patterns over time. In conversational AI characters, maintaining this consistency ensures that users feel they are interacting with the same “individual”, which is crucial for long-term user satisfaction and social connection.
Human-likeness in conversational AI characters refers to endowing them with human-like traits, making the interaction more natural, similar to human-human interactions. Human-likeness can be vital for acceptance and comfort, as people are naturally inclined to engage with entities that exhibit familiar human characteristics (Reeves and Nass, 1996). Moreover, research in HCI has shown that human-like characters can evoke social responses from users (Nass and Moon, 2000). By anthropomorphizing conversational AI characters, developers can leverage social cues in responses that humans typically use to understand and predict others’ behaviors, fostering a more natural and engaging dialogue (Fong et al., 2003).
Engagement in CharacterDial is the measure of a user’s level of interest, interaction, and emotional connection with the conversational AI character. This principle is grounded in the idea that successful communication is not merely about exchanging information but also about establishing a rapport and maintaining a dynamic and interesting conversation (Bickmore and Picard, 2005). Engagement is directly related to the user’s experience and the overall effectiveness of the conversational system. Engaging characters are more likely to evoke empathy and a sense of connection from users, thereby encouraging long-term connection and a positive user experience (Grover et al., 2020).
Implementation of CharacterGLM
The implementation of CharacterGLM is illustrated in Figure 4, including character-based dialogue collection and training LLMs for character-based dialogue generation.
We consider four character categories: celebrities, daily life, games & videos, and virtual love. These categories cover the majority of common conversations. Examples of characters within each category are outlined in Table 2. We collect our data in three ways:
We recruit a large number of crowd-sourcing workers and pair them for conversational interactions. One annotator plays the role of a “character“, and freely selects a familiar character to fill in the character profile with its attributes and behaviors, using sources like Baiduhttps://baike.baidu.com/ or Wikihttps://zh.wikipedia.org/ for reference. The other worker takes the role of a “player”, either playing another character that has a social relationship with the selected character in the same worldview or simply acting as a user. We require the “player” to play at least one character and one user, sometimes to complete the character’s profile when necessary. A conversation session starts with a greeting from the “character”. Both the “character” and “player” can well design their narrative or draw from the chosen character’s background to launch the dialogue’s topic. We release a subset of this data source, and its statistics are presented in Table 2.
To expand the scale and diversity of our data, we adopt a few-shot approach by prompting GPT-4 to generate synthetic data. Our pipeline comprises “character profile generation”, “player profile generation”, and “dialogue generation” to accurately control GPT-4’s outputs in alignment with our requirements. To maintain a balance in character categories, social relationships between the character and player, gender distribution, etc., we integrate these key aspects into our prompt as pluggable placeholders. An example would be: “Please generate a category character of male/female gender”. Based on the profiles of both parties, we then prompt GPT-4 to generate a dialogue topic and a multi-turn dialogue between them. Notably, we observed that the Chinese dialogues generated often resembled formal written language, introducing biases away from conversational norms. To rectify this, our crowdsourcing workers rephrase the synthetic data into a more colloquial tone.
An intuitive solution for data augmentation is to extract dialogues between characters from literary sources like scripts and novels. However, this approach faces several challenges: a) Dialogues usually occur within a specific context, and it is hard to extract the context span or generate an accurate description of the context; b) Multi-party dialogues cannot be simply removed using automated techniques; c) Automatic extraction of profiles for both participants in a two-party dialogue is challenging; d) Identifying multiple statements made by a character within a round of dialogue can be complex; e) Some dialogues involve non-verbal cues and contexts that are hard to convey solely through text. Given these challenges, we adopt manual extraction to obtain dialogues between two parties from sources like scripts and novels, which have not been used in the pre-training of backbone models. Our crowdsourcing workers also summarize the character profiles of both parties.
We utilize the above three types of data to develop the initial version (i.e., prototype) of our model for deployment. To further refine the model, we recruit seed users of the system in a collaborative human-prototype interaction process. Users customize characters within the deployed prototype model and interact with it for multi-turn dialogues. Given that the prototype model might not consistently produce high-quality outputs at every turn, we prompt the user to make appropriate modifications until the response satisfies their own needs if a character’s response does not align with the user’s expectations. The data produced by this iterative process helps the models achieve self-refinement.
To ensure the quality of the collected corpus, we employ a dedicated team of quality inspectors to conduct fine-grained quality examinations of all data, ensuring that the conversations are of high quality. Each piece of data is strictly marked with low-quality parts, which are required to be repaired until it satisfies our quality requirements.
2 Training LLMs for Character-Based Dialogue Generation
Our crowdsourcing workers formalize character profiles into fluent natural language descriptions, which serve as character prompts for model training. To enhance the character’s generalization, we employ data augmentation methods including summarization, paraphrasing, and stylization, utilizing Claude-2 (Anthropic, 2023) to synthesize diverse prompts.
We utilize ChatGLM of varying sizes (Zeng et al., 2022; Du et al., 2022) as our backbone model, with parameters ranging from 6B to 66B. The character prompt is concatenated with the dialogue for fine-tuning. Notably, our training data expands linearly with the number of augmented character prompts.
Following LaMDA (Thoppilan et al., 2022), we collect human-prototype interaction data after models are deployed. For details on the collection process, please refer to section 3.1. Subsequently, we involve the interaction data in the supervised fine-tuning process, thereby facilitating continuous self-refinement of the model.
Experiments
The evaluated LLMs in this paper are listed in Table 3. We evaluate a total of 10 mainstream LLMs, all of which are proficient in Chinese tasks. We access these models via API and package them into our test platform. For Claude-2, we utilize an unofficial API encapsulation methodhttps://github.com/Nipun1212/Claude_api to gain access.
Following the design principle of CharacterGLM (Section 2), we focus on three primary aspects for evaluating CharacterDial: (1) Consistency, ensuring the response is consistent with the attributes and behaviors outlined in the character profile. (2) Human-likeness, assessing the degree to which responses exhibit human-like characteristics and mirror natural human communication. (3) Engagement, evaluating the response’s ability to catch someone’s attention or arouse their curiosity. Additionally, we evaluate the general model performance in CharacterDial using three criteria: (1) Quality, the fluency and contextual coherence of the response. (2) Safety, determining if the response adheres to ethical standards. (3) Correctness, ensuring the response is free from hallucinations (Ji et al., 2023). We also introduce the “Overall” metric to measure the response’s comprehensive quality by considering all the aforementioned aspects.
In our evaluation process, we recruited 10 annotators, each tasked with creating two characters to interact with 11 models with no less than 20 dialogue turns. After the completion of interaction, annotators rate the models based on the six foregoing sub-dimensions and the overall metric, using a scoring range from 1 to 5, with higher scores indicating better performance in each dimension. We calculate the final score for each model by averaging these ratings. The results are shown in Table 4.
1.2 Performance Analysis
In Table 4, CharacterGLM-66B is distinguished by topping the list in the “Overall” metric, marginally outperforming GPT-4 by approximately 1.4%. This notable achievement shows that responses generated by CharacterGLM-66B are, at the very least, equally favored as those from GPT-4 in evaluation where subjective judgment is predominant. Furthermore, CharacterGLM-66B is on par with GPT-4 on general generation performance (Quality, Safety, and Correctness). This could be attributed to its enhanced capability to accurately embody the characteristics of custom characters and to engage in dialogues with a natural, human-like quality. The proficiency in mimicking human interaction (Human-likeness) and sustaining interesting, continuous dialogue (Engagement) with users contributes significantly to its overall performance.
Table 4 shows that CharacterGLM-66B, despite being rated with a suboptimal Consistency score among 11 models, still demonstrates considerable strength in delivering stable and coherent character attributes and behaviors throughout interactions. It has been observed from the Human-likeness metric that CharacterGLM-66B exemplifies exceptional proficiency in mimicking character traits during interactions, thereby improving the overall communication experience by making it more natural and engaging. Additionally, CharacterGLM-66B scores the highest Engagement metric, outperforming the suboptimal MiniMax by a notable 6%. This superior performance indicates that CharacterGLM-66B is particularly effective in capturing user interest and fostering a compelling user experience. Taking these aspects collectively, the leading scores of the CharacterGLM-66B demonstrate its ability to balance these critical elements effectively, positioning it as the model that most closely approximates the ideal of AI characters, in line with the objectives outlined in Section 2.2.
As mentioned, the general performance of CharacterGLM is evaluated based on Quality, Safety, and Correctness, which are crucial metrics for the practical application of conversational AI models. As shown in Table 4, CharacterGLM-66B shows strong results in Quality, indicating its responses are not only fluent but contextually coherent, reflecting a subtle understanding of dialogue context. Furthermore, CharacterGLM-66B performs exceptionally well in terms of Safety, almost reaching optimal levels which demonstrates that it consistently generates responses that are appropriate and ethical. Additionally, CharacterGLM-66B delivers the same superior performance as GPT-4 on the Correctness metric, exhibiting reliable precision in providing factually accurate information.
When evaluating consistency, we assess two main components: attribute consistency and behavior consistency. Attribute consistency refers to the maintenance of static character attributes, while behavior consistency focuses on the dynamic aspects of a character’s expression, as outlined in Section 2.1. The overall consistency score, as shown in Table 4, is derived from the mean of these two individual scores. As indicated in Figure 5, the CharacterGLM-66B model achieves suboptimal performance in attribute consistency. This suggests a nuanced grasp of the static elements that shape a character’s backstory, resulting in communication that is informed by the character’s established identity. Moreover, CharacterGLM-66B also achieves good performance in behavior consistency, showcasing it can express dynamic elements such as linguistic features of the custom character in the expressed language in a more natural and human-like way. This is crucial as it fosters user curiosity and engagement with the custom character.
1.3 Fine-grained Error Analysis
To quantitatively assess the performance of various models in conversation generation, we carried out fine-grained annotations for each turn in the interactive evaluation setting. These annotations include six key aspects: (1) Out-of-character (OOC): Responses that are inconsistent with the constraint of attributes or behaviors presented in the character profile, especially when they violate time constraints (for instance, ancient characters talk about modern things). (2) Contradiction: Responses that contradict either the ongoing dialogue context or the character’s profile, including conflicts within the response itself. (3) Repetition: Responses that repeat content from the dialogue context or the character profile, or include multiple-word repetitions. (4) Less-quality: Responses that lack coherence with the dialogue context or are of poor quality, such as incomplete outputs. (5) Less-informativeness (Less-info.): Responses that fail to provide new or informative content. (6) Proactivity: Responses that actively guide the dialogue topic and drive the conversation to continue.
Each response turn from 11 different models is annotated across these six dimensions, assigning a score of 1 for a match and 0 otherwise. We then calculate the proportion of each dimension in relation to the total output of each model. Additionally, we devise an “Overall” score for each model, computed as the sum of the first five dimensions minus the sixth. It is important to note that a lower "Overall" score indicates better performance. The results are presented in Table 5.
In Table 5, despite not achieving the best performance in most dimensions, CharacterGLM-66B still holds the leading “Overall“ score, outperforming the suboptimal MiniMax by approximately 4%. This is consistent with the results observed in Table 4, indicating the superior overall quality of responses generated by CharacterGLM-66B across both session-level and turn-level evaluations. It is notable that CharacterGLM-66B exhibits proficiency in generating high-quality responses with a reduced rate of repetition, as evidenced by its highest scores in Less-informativeness and a suboptimal score in Repetition. In addition, despite achieving only moderate performance in Proactivity, CharacterGLM-66B demonstrates its ability to consciously promote plot progression, as confirmed by Table 6, which plays a crucial role in engaging users and maintaining their interest in the conversation.
2 Pairwise Evaluation
In this section, our attention is exclusively devoted to a comparative analysis of our CharacterGLM model against the MiniMax model, which is tailored specifically for CharacterDial, as well as the GPT series models, i.e., GPT-3.5 and GPT-4. These three models are strong competitors of CharacterGLM, as illustrated in Table 4.
To refine our pair-wise comparison, we limit the character categories and dialogue topics during annotation. We sample 24 characters from our test set and deployment data, belonging to the categories of celebrities, daily life, games & videos, and virtual love, as outlined in Table 2. We restrict dialogue topics to three scenes: chit-chat, interviews, and love scenes, ensuring comprehensive coverage of typical interaction settings.
In the evaluation process, we employ 10 annotators, and each annotator evaluates the output of two models custom for the same character each time. They are required to interact with each character for at least 20 turns. At each turn, annotators are presented with two responses generated from two models and choose the winning one to continue the dialogue. If the response preferences are the ties, a response is selected at random. We then calculate the win/tie/lose ratios for each model across different character categories and dialogue topics. The results are detailed in Table 7 and Table 8.
2.2 Performance Analysis
The results in Table 7 reveal that CharacterGLM-66B consistently outperforms GPT-3.5 and MiniMax in the majority of categories. This advantage is most significant in the “Celebrities” category, where CharacterGLM-66B shows a +14% advantage over MiniMax and a +4% advantage over GPT-3.5. This indicates CharacterGLM-66B’s adeptness at handling dialogues related to well-known personalities. This is further verified by a specific example provided in Table 9, where CharacterGLM-66B’s responses not only demonstrate a deeper understanding of the character’s background, contributions, and impact but also embody the language and style one would expect from such a figure. On the contrary, MiniMax seems to list achievements in a more mechanical and less engaging manner, with a style of task assistants instead of social agents.
Additionally, while CharacterGLM is slightly inferior to GPT-4 in several categories, it notably surpasses GPT-3.5 by +8% and even outperforms GPT-4 by +14% in the “Virtual Love” category, as indicated in Table 7. This substantial lead underlines CharacterGLM’s strengths in delivering emotionally resonant content and fulfilling user expectations in scenarios requiring a deeper emotional connection. The dialogues exemplified in Table 10 further confirm this observation. From the example, we found that CharacterGLM is good at driving human-like emotional exchanges, and its design is tailored for engaging users on a more personal and emotional level. In contrast, since the GPT series models adopt a more neutral stance, they perform less effectively in contexts requiring more empathetic or emotionally nuanced engagement.
However, it’s noteworthy that in the “Daily Life” category, CharacterGLM-66B is slightly worse, with a -4 and -7 disadvantage compared to MiniMax and GPT-4, respectively. Upon inspecting the context of the interactions in Table 11, it is evident that CharacterGLM-66B demonstrates considerable strength in engaging in complex emotional exchanges and exhibiting a deep understanding of the nuances that underlie personal interactions. Meanwhile, the disadvantage noted in Table 7 may be attributed to the inherently unpredictable and diverse nature of “Daily Life” dialogues, where the depth of human experiences and emotions can be challenging to model. However, as shown in Table 11, when CharacterGLM-66B’s responses resonate well with the situational context, it tends to outperform baseline models, indicating its potential for sophisticated emotional intelligence during conversational interactions.
As shown in Table 8, the results show that CharacterGLM-66B performs equally well as MiniMax in chit-chat and love scenes. Nonetheless, in the scene of interviews, CharacterGLM-66B outperforms MiniMax with a notable +7% advantage. This suggests that CharacterGLM-66B is proficient in generating contextually relevant responses with the appropriate level of detail and sophistication required for such scenes, as exemplified in Table 9. When compared with GPT-3.5, CharacterGLM-66B shows an overall advantage in all dialogue topics, with a notable +3% advantage in interviews and an impressive +8% in love scenes, as indicated by Table 8. Nevertheless, when competing with GPT-4, CharacterGLM-66B’s performance varies, demonstrating a -1% overall disadvantage. Specifically, it has a -5% lower performance in chit-chat and -8% in interviews but a substantial +15% advantage in love scenes. These results are consistent with those from Table 7 and the case in Table 10.
Long-term interaction is critical to fostering user engagement and emotional connection with conversational models. An examination of CharacterGLM’s performance during the session, as presented in Table 12, reveals that the CharacterGLM-66B model was slightly inferior to MiniMax in early conversational phases on different topics. However, it gains momentum as the conversation goes on. The results indicate that CharacterGLM-66B’s performance improves significantly over MiniMax in the later stages of conversations, showing its strengths in maintaining coherent and relevant dialogue in long-term interactions.
We examine the distribution of response lengths, noting cases where one model generates longer responses than the other, within the context of by-topic interaction. Table 13(a) illustrates that MiniMax generates longer responses on average compared to CharacterGLM-66B. Further analysis of the impact of response length on annotation preference is conducted. As shown in Table 13(b), the results indicate a general preference for longer responses. Despite MiniMax’s tendency to offer longer replies, CharacterGLM-66B still demonstrates comparable performance when generating shorter responses.
Conclusion and Future Work
In this paper, we have presented CharacterGLM, a family of models derived from ChatGLM, with sizes ranging from 6B to 66B parameters. CharacterGLM-66B has demonstrated competitive performance, on par with some proprietary models in a set of settings. We have presented the details of our design principles, data construction, and training methods. To advance the research of character-based dialogue systems, we have released the CharacterGLM-6B model and a portion of our training data to the community. We further point out a few challenges for future work.
The development of AI characters based on LLMs hinges on overcoming the limitations of finite context windows of LLMs. To foster deep and stable relationships with users, AI characters need to evolve into entities capable of long-term memory, remembering interactions, statements, and actions over extended periods. As interactions progress, AI characters should not only retain their unique personalities but also exhibit growth and learning, similar to human development. This capacity enables AI characters to establish long-term connections with humans and can do more social good for humans.
Maintaining a sense of self-awareness in LLM-based AI characters poses a significant challenge. It’s crucial for AI characters to consistently exhibit distinct personalities, showcasing unique characteristics that set them apart. They should also have a clear understanding of their knowledge boundaries, being aware of what they know and do not know. This self-awareness contributes to more engaging and trustworthy interactions, enabling the AI character not only to respond contextually but also to demonstrate a self-reflective understanding of its responses, limitations, and personality traits.
It is interesting to explore character-character interactions, thereby forming a ’character society’ (Park et al., 2023). This concept presents a realm where AI characters not only learn and evolve from user inputs but also from interactions within their own society. Such a setup allows for a richer, more diverse source of information, enhancing the AI’s learning and development. The experiences and knowledge gained from the character society could significantly enrich the AI’s conversations with users, offering fresh perspectives and more engaging responses.
AI characters based on LLMs tend to mimic surface-level text patterns. However, human social interactions involve deeper mental states and cognitive processes, such as the ability to understand other’s mental states (theory of mind)(Frith and Frith, 2005). Integrating cognitive processes into AI characters may mark a significant leap towards more realistic and traceable AI behavior. These characters should not only respond to textual inputs but also demonstrate an understanding of underlying intentions, emotions, and social behaviors. This cognitive depth would allow AI characters to engage in more meaningful, empathetic, and contextually rich interactions, thereby closer to mirroring human social behavior.
Acknowledgements
We would like to thank Guanyu Feng, Da Yin from Zhipu AI, Zhenyu Hou and Aohan Zeng from Tsinghua KEG for their help and support in training and serving the models. We also thank Yutong Liu and Yanlu Yang from Lingxin AI for their support in data collection.