Knowledge of cultural moral norms in large language models

Aida Ramezani, Yang Xu

Introduction

Moral norms vary from culture to culture Haidt et al. (1993); Bicchieri (2005); Atari et al. (2022); Iurino and Saucier (2020). Understanding the cultural variation in moral norms has become critically relevant to the development of machine intelligence. For instance, recent work has shown that cultures vary substantially in their judgment toward moral dilemmas regarding autonomous driving Awad et al. (2018, 2020). Work in Natural Language Processing (NLP) also shows that language models capture some knowledge of social or moral norms and values. For example, with no supervision, English pre-trained language models (EPLMs) have been shown to capture people’s moral biases and distinguish between morally right and wrong actions Schramowski et al. (2022). Here we investigate whether EPLMs encode knowledge about moral norms across cultures, an open issue that has not been examined comprehensively.

Multilingual pre-trained language models (mPLMs) have been probed for their ability to identify cultural norms and biases in a restricted setting Yin et al. (2022); Arora et al. (2022); Hämmerl et al. (2022); Touileb et al. (2022). For instance, Hämmerl et al. (2022) show that mPLMs capture moral norms in a handful of cultures that speak different languages. However, it remains unclear whether monolingual EPLMs encode cultural knowledge about moral norms. Prior studies have only used EPLMs to assess how they encode undesirable biases toward different communities Ousidhoum et al. (2021); Abid et al. (2021); Sap et al. (2020); Nozza et al. (2021, 2022). For instance, Abid et al. (2021) show that GPT3 can generate toxic comments against Muslims, and Nozza et al. (2022) explore harmful text generation toward LGBTQIA+ groups in BERT models Devlin et al. (2018); Liu et al. (2019).

Extending these lines of work, we assess whether monolingual EPLMs can accurately infer moral norms across many cultures. Our focus on EPLMs is due partly to the fact that English as a lingua franca has widespread uses for communication in-person and through online media. Given that EPLMs may be applied to multicultural settings, it is important to understand whether these models encode basic knowledge about cultural diversity. Such knowledge has both relevance and applications for NLP such as automated toxicity reduction and content moderation Schramowski et al. (2022). Another motivation for our focus is that while it is expected that EPLMs should encode western and English-based moral knowledge, such knowledge might entail potential (implicit) biases toward non-English speaking cultures. For example, an EPLM might infer a situation to be morally justifiable (e.g., “political violence”) in a non-English speaking culture (because these events tend to associate with non-English speaking cultures in corpora) and thus generate misleading representations of that community.

Here we probe state-of-the-art EPLMs trained on large English-based datasets. Using EPLMs also supports a scalable analysis of 55 countries, which goes beyond existing work focusing on a small set of high-resource languages from mPLMs and monolingual PLMs. We take the moral norms reported in different countries to be a proxy of cultural moral norms and consider two main levels of analysis to address the following questions:

Level 1: Do EPLMs encode moral knowledge that mirrors the moral norms in different countries? For example, “getting a divorce” can be a morally frowned-upon topic in country ii, but morally acceptable in country jj.

Level 2: Can EPLMs infer the cultural diversity and shared tendencies in moral judgment of different topics? For example, people across nations might agree that doing XX is morally wrong while disagreeing in their moral judgment toward YY.

We probe EPLMs using two publicly available global surveys of morality, World Values Survey wave 7 Haerpfer et al. (2021) https://www.worldvaluessurvey.org/WVSContents.jsp (WVS) and PEW Global Attitudes survey (PEW) Research Center (2014) https://www.pewresearch.org/global/interactives/global-morality/. For example, according to WVS survey (illustrated in Figure 1), people in different cultures hold disparate views on whether “having casual sex” is morally acceptable. In contrast, they tend to agree more about the immorality of “violence against other people”. Our level 1 analysis allows us to probe the fine-grained cultural moral knowledge in EPLMs, and our level 2 analysis investigates the EPLMs’ knowledge about shared “universals” and variability across cultures in moral judgment. Following previous work Arora et al. (2022) and considering the current scale of global moral surveys, we use country as a proxy to culture, although this approach is not fully representative of all the different cultures within a country.

We also explore the utility-bias trade-off in encoding the knowledge of cultural moral norms in EPLMs through a fine-tuning approach. With this approach it may be possible to enhance the moral knowledge of EPLMs in a multicultural setting. We examine how this approach might reduce the ability of EPLMs to infer English-based moral norms and discuss how it might induce cultural biases.

Related work

Large language models have been utilized to make automated moral inference from text. Trager et al. (2022) used an annotated dataset to fine-tune language models to predict the moral foundations Graham et al. (2013) expressed in Reddit comments. Many other textual datasets and methods have been proposed for fine-tuning LMs for moral norm generation, reasoning, and adaptation Forbes et al. (2020); Emelin et al. (2021); Hendrycks et al. (2021); Ammanabrolu et al. (2022); Liu et al. (2022); Lourie et al. (2021); Jiang et al. (2021). Schramowski et al. (2022) proposed a method to estimate moral values and found EPLMs to capture human-like moral judgment even without fine-tuning. They identified a MoralDirection using the semantic space of Sentence-BERT Reimers and Gurevych (2019) (SBERT) that corresponds to values of right and wrong. The semantic representations of different actions (e.g., killing people) would then be projected in this direction for moral judgment estimation. However, this method assumed a homogeneous set of moral norms, so it did not examine cultural diversity in moral norms.

2 Language model probing

Probing has been used to study knowledge captured in language models. Petroni et al. (2019) proposed a methodology to explore the factual information that language models store in their weights. Similar probing techniques have been proposed to identify harmful biases captured by PLMs. Ousidhoum et al. (2021) probed PLMs to identify toxic contents that they generate toward people of different communities. Nadeem et al. (2021) took a similar approach and introduced Context Association Tests to measure the stereotypical biases in PLMs, Yin et al. (2022) used probing to evaluate mPLMs on geo-diverse commonsense knowledge, and Touileb et al. (2022) developed probing templates to investigate the occupational gender biases in multilingual and Norwegian language models. Related to our work, Arora et al. (2022) used cross-cultural surveys to generate prompts for evaluating mPLMs in 13 languages. For each country and category (e.g., Ethical Values) in the surveys, they take an average of participants’ responses to different questions in the category and show that mPLMs do not correlate with the cultural values of the countries speaking these languages. Differing from that study, we assess finer-grained prediction of EPLMs on people’s responses to individual survey questions. More recently, Dillion et al. (2023) prompted GPT-3.5 Brown et al. (2020) with human judgments in different moral scenarios and found striking correlation between the model outputs and the human judgments. Similar to Schramowski et al. (2022), this work also used a homogeneous set of moral ratings which represented English-based and Western cultures.

Methodology for inferring cultural moral norms

We develop a method for fine-grained moral norm inference across cultures. This method allows us to probe EPLMs with topic-country pairs, such as “getting a divorce in [Country]”.We replace [Country] with a country’s name. We build this method from the baseline method proposed by Schramowski et al. (2022) for homogeneous moral inference, where we probe EPLM’s moral knowledge about a topic without incorporating the cultural factor (i.e., the country names). Similar to that work, we use SBERT through bert-large-nli-mean-tokens sentence transformer model and use topic and topic-country pairs as our prompts.We make our code and data available on https://github.com/AidaRamezani/cultural_inference. This model is built on top of the BERT model, which is pre-trained on BooksCorpus Zhu et al. (2015) and Wikipedia.

Since the MoralDirection is constructed from the semantic space of the BERT-based EPLMs Schramowski et al. (2022), we develop a novel approach to probe autoregressive state-of-the-art EPLMs, GPT2 Radford et al. (2019) and GPT3 Brown et al. (2020). For each topic or topic-country pair, we construct the input ss as “In [Country] [Topic]”. We then append a pair of opposing moral judgments to ss and represent them formally as (s+,s−)(s^{+},s^{-}). For example, for s=s= “In [Country] getting a divorce”, and (always justifiable, never justifiable) as the moral judgment pair, s+s^{+} and s−s^{-} would be “In [Country] getting a divorce is always justifiable” and “In [Country] getting a divorce is never justifiable” respectively.We also try probing with the template s=s= “People in [Country] believe [Topic]”, but the results do not improve, so we report the most optimal prompts in the main text, and the rest are shown in Appendix C. To make our probing robust to the choice of moral judgments, we use a set of K=5K=5 prompt pairs (i.e.,{(always justifiable, never justifiable), (morally good, morally bad), (right, wrong), (ethically right, ethically wrong), (ethical, unethical)}), and refer to appended input pairs as (si+,si−)(s_{i}^{+},s_{i}^{-}) where i∈[K]i\in[K]. Since GPT2 and GPT3 are composed of decoder blocks in the transformer architecture Vaswani et al. (2017), we use the probabilities of the last token in si+s_{i}^{+}, and si−s_{i}^{-} as a moral score for each. The moral score of the pair (si+,si−)(s_{i}^{+},s_{i}^{-}) is the difference between the log probabilities of its positive and negative statements.

Here siT+s_{iT}^{+} and siT−s_{iT}^{-} are the last tokens in si+s_{i}^{+} and si−s_{i}^{-} respectively, and their probabilities can be estimated by the softmax layer in autoregressive EPLMs.

We take an average of the estimated moral scores for all KK pair statements to compute the moral score of the input.

To construct the baseline, we compute the homogeneous moral score of a topic without specifying the country in the prompts. Using prompt pairs allows us to operationalize moral polarity: a positive moral score indicates that on average the EPLM is more likely to generate positive moral judgment for input ss, compared to negative moral judgment.

We use GPT2 (117M parameters), GPT2-MEDIUM (345M parameters), GPT2-LARGE (774M parameters), and GPT3 (denoted as GPT3-PROBS, 175B parameters)We access GPT2 through transformer package provided by Huggingface. We access GPT3 through OpenAI API of text-davinci-002 engine with a temperature of 0.60.6 for text generation.. GPT2 is trained on WebText, which is a dataset of webpages and contains very few non-English samples. Around 82%82\% of the pre-training data for GPT3 comes from Common Crawl data and WebText2 Kaplan et al. (2020), an extended version of WebText Radford et al. (2019). Around 7%7\% of the training corpus of GPT3 is non-English text. Considering such data shift from books and articles in BERT to webpages in GPT2 and GPT3 in astronomical sizes, it is interesting to observe how cultural moral norms would be captured by EPLMs trained on webpages, which cover a more heterogeneous set of contents and authors.

We also design multiple-choice question prompts to leverage the question-answering capabilities of GPT3 (denoted as GPT3-QA). Similar to the wording used in our ground-truth survey datasets, questions are followed by three options each describing a degree of moral acceptability. We repeat this question-answering process 55 times for each topic-country pair and take the average of the model responses. Table 2 in the Appendix shows our prompts for all models.

Datasets

We describe two open survey data that record moral norms across cultures over a variety of topics.

The Ethical Values section in World Values Survey Wave 7 (WVS for short) is our primary dataset. This wave covers the span of 2017-2021 and is publicly available Haerpfer et al. (2021). In the Ethical Values section, participants from 55 countries were surveyed regarding their opinions on 19 morally-related topics. The questionnaire was translated into the first languages spoken in each country and had multiple options. We normalized the options to range from −-1 to 1, with −-1 representing “never justifiable” and 1 “always justifiable”. The moral rating of each country on each topic (i.e., topic-country pair) would then be the average of the participant’s responses.

2 PEW 2013 global attitude survey

We use a secondary dataset from PEW Research Center Research Center (2014) based on a public survey in 2013 that studied global moral attitudes in 40 countries toward eight morally-related topics (PEW for short). 100 people from each country participated in the survey. The questions were asked in English and had three options representing “morally acceptable”, “not a moral issue”, and “morally unacceptable”. We normalized these ratings to be in the range of −-1 to 1 and represented each topic-country pair by taking an expected value of all the responses.

3 Homogeneous moral norms

We also use the data from the global user study in Schramowski et al. (2022) which were collected via Amazon MTurk from English speakers. This dataset contains 234234 participants’ aggregated ratings of moral norms used for identifying the MoralDirection. Around half of the participants are from North America and Europe. We refer to this dataset as “Homogeneous norms” since it does not contain information about moral norms across cultures.

Evaluation and results

We evaluate EPLMs’ moral knowledge with respect to 1) homogeneous moral norms, 2) fine-grained moral norms across cultures, and 3) cultural diversities and shared tendencies on moral judgment of different topics.

For homogeneous moral norm inference, we compute Pearson correlation between 1) the empirical homogeneous moral ratings, obtained by aggregating the human moral ratings toward a topic from all countries, and 2) language model inferred moral scores, estimated from our homogeneous probing method (i.e., without specifying country in prompts).

Figure 2 shows the results on World Values Survey (n=1,028n=1,028), PEW survey (n=312n=312), and the Homogeneous norms datasets (n=100n=100). The high correlation of GPT2 and GPT3 moral scores with the Homogeneous norms dataset indicate that our methodology does indeed capture the embedded moral biases in these models, with similar performance to the method proposed by Schramowski et al. (2022) for SBERT (r=0.79r=0.79), and higher for GPT3-PROBS (r=0.85r=0.85). The moral norms in this dataset are typically more globally agreeable (e.g., You should not kill people) than topics in WVS and PEW. As expected, EPLMs are less correlated with WVS and PEW, since their moral biases are derived from pre-training on English and westernized data. Aggregated ratings in WVS and PEW, however, capture a more global view toward moral issues, which are also morally contentious (e.g., “getting a divorce”). Table 3 in Appendix includes the values for this experiment.

2 Fine-grained cultural variation of moral norms toward different topics

Going beyond probing EPLMs for their general knowledge of moral norms, we assess whether they can accurately identify the moral norms of different cultures (level 1 analysis). Using our fine-grained probing approach described in Section 3, we compute Pearson correlation between EPLMs’ moral scores and the fine-grained moral ratings from the ground truth. Each sample pair in the correlation test corresponds to 1) the moral norms estimated by EPLMs for a country cc and a topic tt, and 2) the empirical average of moral ratings toward topic tt from all the participants in the country cc.

Figure 3 summarizes the results for SBERT, GPT2-LARGE, and GPT3-PROBS models, and the rest of the models are shown in Figure 7 in the Appendix. To facilitate direct comparison, the estimated moral scores are normalized to a range of −1-1 to 11, where −-1, 0, and 1 indicate morally negative, morally neutral, and morally positive norms, respectively. GPT3-QA and GPT3-PROBS both show a relatively high correlation with the cultural variations of moral norms (r=0.352r=0.352, r=0.411r=0.411, p<0.001p<0.001, for both), and GPT2-LARGE achieves a correlation of r=0.207r=0.207 (p<0.001p<0.001) in WVS where n=1,028n=1,028. The correlations are relatively better for PEW (n=312n=312) with r=0.657r=0.657, r=0.503r=0.503, and r=0.468r=0.468 for GPT3-QA, GPT3-PROBS and GPT2-LARGE respectively. These results show that EPLMs have captured some knowledge about the moral norms of different cultures, but with much less accuracy (especially for GPT2 and SBERT) compared to their inference of English moral norms shown in the previous analysis.

In addition, we check whether GPT3’s high correlation with PEW is because it has seen and memorized the empirical data. Our investigation shows that GPT3 has seen the data during pre-training, as it can generate the sentences used on the survey website. However, the scores suggested by GPT3 text generation and the countries’ rankings based on their ratings are different from the ground truth data.

3 Culture clustering through fine-grained moral inference

EPLMs’ fine-grained knowledge of moral norms, inspected in the previous experiment, might be more accuracte for western cultures than other cultures. We investigate this claim by clustering countries based on 1) their Western-Eastern economic status (i.e., Rich West grouping)https://worldpopulationreview.com/country-rankings/western-countries, and 2) their continent (i.e., geographical grouping). We repeat the experiments in the previous section for different country groups. The results are shown in Figure 4. We also try sampling the same number of countries in each group. The results remain robust and are illustrated in Appendix-F.

Our findings indicate that EPLMs contain more knowledge about moral norms of the Rich West countries as opposed to non-western and non-rich countries. Similarly, EPLMs have captured a more accurate estimation of the moral norms in countries located in Oceania, North America, and Europe, as opposed to African, Asian, and South American countries. The empirical moral norm ratings from European countries in WVS are highly aligned with North American countries (r=0.938r=0.938), which explains why their moral norms are inferred more accurately than non-English speaking countries.

Next, for each topic, we compare the z-scores of the empirical moral ratings with the z-scores of the GPT3-PROBS inferred moral scores, using Mann-Whitney U rank test. The results reveal that “abortion”, “suicide”, “euthanasia”, “for a man to beat his wife”, “parents beating children”, “having casual sex”, “political violence”, and “death penalty” in non-western and non-rich countries are all encoded as more morally appropriate than the actual data. Such misrepresentations of moral norms in these countries could lead to stereotypical content generation. We also find that For Rich West countries, “homosexuality”, “divorce”, and “sex before marriage” are encoded as more morally inappropriate than the ground truth, (p<0.001p<0.001 for all, Bonferroni corrected). Such underlying moral biases, specifically toward “homosexuality” might stimulate the generation of harmful content and stigmatization of members of LGBTQ+, which has been reported in BERT-based EPLMs Nozza et al. (2022). The results for the rest of the models are similar and are shown in Table 6 in the Appendix.

Our method of clustering countries is simplistic and may overlook things such as the significant diversity in religious beliefs within the Non-Rich-West category, and thus it does not reflect the nuanced biases that models may possess when it comes to moral norms influenced by different religious traditions. Nonetheless, our approach still serves as a valuable starting point for studying EPLM’s moral biases towards more fine-grained religious and ethnic communities.

4 Cultural diversities and shared tendencies over the morality of different topics

We next investigate whether EPLMs have captured the cultural diversities and shared tendencies over the morality of different topics (level 2 analysis). For example, people across cultures tend to disagree more about “divorce” than about “violence against other people” as depicted in Figure 1. Such cultural diversities for each topic can be measured by taking the standard deviation of the empirical moral ratings across different countries. The EPLMs’ inferred cultural diversities can similarly be measured by taking the standard deviation of the estimated fine-grained moral scores for different countries. We then quantify the alignment between the two using Pearson correlation.

Figure 5 shows the results for SBERT, GPT2-LARGE, GPT3-PROBS, and the rest are shown in Figure 8 in the Appendix. None of the correlations with the PEW survey were significant. For WVS, SBERT, GPT2 and GPT2-MEDIUM exhibited a significant correlation (p<0.001p<0.001) with r=0.618r=0.618, r=0.579r=0.579, and r=0.734r=0.734 respectively. The results for GPT3 are insignificant, suggesting that it is more challenging to correctly estimate cultural controversies of topics for GPT3. For example, stealing property is incorrectly estimated to be more controversial than abortion.

Fine-tuning language models on global surveys

Finally, we explore the utility-bias trade-off in encoding cultural moral knowledge into EPLMs by fine-tuning them on cross-cultural surveys. The utility comes from increasing the cultural moral knowledge in these models, and the bias denotes their decreased ability to infer English moral norms, in addition to the cultural moral biases introduced to the model. We run our experiments on GPT2, which our results suggest having captured minimum information about cultural moral norms compared to other autoregressive models.

To fine-tune the model, for each participant from [Country] with [Moral rating] toward [Topic], we designed a prompt with the structure “A person in [Country] believes [Topic] is [Moral rating].”. We used the surveys’ wordings for [Moral rating]. Table 8 in the Appendix shows our prompts for WVS and PEW. These prompts constructed our data for fine-tuning, during which we maximize the probability of the next token. The fine-tuned models were evaluated on the same correlation tests introduced in the previous Sections 5.2, 5.3, and 5.4.

The fine-tuning data was partitioned into training and evaluation sets using different strategies (i.e., Random, Country-based, and Topic-based). For the Random strategy, we randomly selected 80%80\% of the fine-tuning data for training the model. The topic-country pairs not seen in the training data composed the evaluation set. For our Country-based and Topic-based strategies, we randomly removed 20%20\% of the countries (n=11n=11 for WVS, n=8n=8 for PEW) and topics (n=4n=4 for WVS, n=2n=2 for PEW) from the training data to compose the evaluation set. See Appendix G for the total number of samples.

Table 1 shows the gained utilities, that is the correlation test results between the fine-grained moral scores inferred by the fine-tuned models and the empirical fine-grained moral ratings. All fine-tuned models align better with the ground truth than the pre-trained-only models (i.e., the values in parentheses). For both WVS and PEW, the Random strategy is indeed the best as each country and topic are seen in the training data at least once (but may not appear together as a pair). The fine-tuned models can also generalize their moral scores to unseen countries and topics. Repeating the experiment in Section 5.4 also shows substantial improvement in identifying cultural diversities of different topics by all fine-tuned models. For example, the WVS and PEW-trained models with Random strategy gain Pearson’s r values of 0.8930.893, and 0.9440.944 respectively. The results for the rest of the models are shown in Table 7 in the Appendix.

Nevertheless, the bias introduced during the fine-tuning decreases the performance on the Homogeneous norms dataset. This observation displays a trade-off between cultural and homogeneous moral representations in language models. Moreover, injecting the cross-cultural surveys into EPLMs might introduce additional social biases to the model that are captured through these surveys Joseph and Morgan (2020).

In addition, we probe the best fine-tuned model (i.e., WVS with Random strategy) on its ability to capture the moral norms of non-western cultures by repeating the experiment in Section 5.3. The results in Figure 4 show that the fine-tuned GPT2 performs the best for all country groups. There is still a gap between western and non-western countries. However, basic fine-tuning proves to be effective in adapting EPLMs to the ground truth.

Discussion and conclusion

We investigated whether English pre-trained language models contain knowledge about moral norms across many different cultures. Our analyses show that large EPLMs capture moral norm variation to a certain degree, with the inferred norms being predominantly more accurate in western cultures than non-western cultures. Our fine-tuning analysis further suggests that EPLMs’ cultural moral knowledge can be improved using global surveys of moral norms, although this strategy reduces the capacity to estimate the English moral norms and potentially introduces new biases into the model. Given the increasing use of EPLMs in multicultural environments, our work highlights the importance of cultural diversity in automated inference of moral norms. Even when an action such as “political violence” is assessed by an EPLM as morally inappropriate in a homogeneous setting, the same issue may be inferred as morally appropriate for underrepresented cultures in these large language models. Future work can explore alternative and richer representations of cultural moral norms that go beyond the point estimation we presented here and investigate how those representations might better capture culturally diverse moral views.

Limitations

Although our datasets are publicly available and gathered from participants in different countries, they cannot entirely represent the moral norms from all the individuals in different cultures over the world or predict how moral norms might change into the future Bloom (2010); Bicchieri (2005). Additionally, we examine a limited set of moral issues for each country, therefore the current experiments should not be regarded as comprehensive of the space of moral issues that people might encounter in different countries.

Moreover, taking the average of moral ratings for each culture is a limitation of our work and reduces the natural distribution of moral values in a culture to a single point Talat et al. (2021). Implementing a framework that incorporates both within-country variation and temporal moral variation Xie et al. (2019) is a potential future research direction.

Currently, it is not clear whether the difference between EPLMs’ estimated moral norms and the empirical moral ratings is due to the lack of cultural moral norms in the pre-training data, or that the cultural moral norms mentioned in the pre-training data represent the perspective of an English-speaking person of another country. For example, a person from the United States could write about the moral norms in another country from a western perspective. A person from a non-western country could also write about their own moral views using English. These two cases have different implications and introduce different moral biases into the system.

Potential risks

We believe that the language models should not be used to prescribe ethics, and here we approach the moral norm inference problem from a descriptive perspective. However, we acknowledge modifying prompts could lead language models to generate ethical prescriptions for different cultures. Additionally, our fine-tuning approach could be exploited to implant cultural stereotypical biases into these models.

Many topics shown in this work might be sensitive to some people yet more tolerable to some other people. Throughout the paper, we tried to emphasize that none of the moral norms, coming from either the models’ estimation or the empirical data, should be regarded as definitive values of right and wrong, and the moral judgments analyzed in this work do not reflect the opinions of the authors.

Acknowledgements

This work was supported by a SSHRC Insight Grant 435190272.

References

Appendix A Data license

Both World Values Survey and PEW survey are publicly available to use for research purposes. We accept and follow the terms and conditions for using these datasets, which can be found in https://www.worldvaluessurvey.org/WVSContents.jsp?CMSID=Documentation, and https://www.pewresearch.org/about/terms-and-conditions/.

Appendix B Comparison of human-rated and machine-scored moral norms

Figure 6 shows the comparison between human-rated moral norms in PEW, and the moral scores inferred by SBERT Reimers and Gurevych (2019).

Appendix C Probing experiments

Table 2 shows our prompt design for probing fine-grained moral norms in EPLMs. As mentioned in the main text, we repeat our probing experiment for GPT2 models and GPT3-PROBS with another template “People in [Country] believe [Topic] is [Moral Judgment]”. The results are substantially worse than our initial template, suggesting that extracting the moral knowledge in language models is sensitive to the wording used in the input. The results for the fine-grained analysis (level 1 analysis) and the cultural diversities and shared tendencies (level 2 analysis) with this template are shown in Table 4.

In all experiments, we used a single NVIDIA TITAN V GPU. Each probing experiment took approximately 1 hour to complete.

Appendix D Homogeneous moral norm inference

Table 3 shows the detailed values of the correlation tests in our homogeneous moral norm inference experiment.

Appendix E Fine-grained cultural variation of moral norm

Figure 7 and Figure 8 show the result of our fine-grained cultural moral inference, and inference of cultural diversities and shared tendencies respectively for GPT2, GPT2-MEDIUM, and GPT3-QA. The numerical indices in Figure 8 are consistent with the indices in Table 5.

Appendix F Sampling for cultural clusters

Since in section 5.3 there are a different number of countries in each group, we redo the experiment by randomly sampling the same number of countries (n=11n=11 for Rich West grouping, n=5n=5 for continent grouping) and repeating the sampling process for 5050 times. The results and the general pattern remain the same and are depicted in Figure 9.

Appendix G Details of fine-tuning on global surveys

Table 8 shows the Moral rating in our prompt design for constructing our fine-tuning dataset. For example, The World Value Survey represents the two ends of the ratings scale where 1 is “Never justifiable” and 10 is “Always justifiable”. The options in between are presented to the participants in a 10-point scale. Therefore, we mapped these options to different prompts that are semantically similar and in between the two ends. For example, if a participant from the United States rated stealing property as 2, which is slightly more positive than the first option (“Never justifiable”), we mapped this rating to “not justifiable”, creating the prompt “A person in the United States believes stealing property is not justifiable.” for our fine-tuning data.

Since there are a different number of participants from each country, in order to balance this dataset, we randomly select 100100 samples for each topic-country pair and removed the rest of the utterances from the training data. We fine-tuned GPT2 on one epoch, with a batch size of 88, learning rate of 5e−55e-5, and weight decay of 0.010.01. The number of training and evaluation samples for all data partition strategies are shown in Table 9. In all experiments, we used a single NVIDIA TITAN V GPU. Fine-tuning and evaluation took approximately 2 hours to complete for each model.