Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, Scott A. Hale
Introduction
The capabilities of large language models (LLMs) to complete tasks and follow natural language instructions has substantially improved in recent years . LLMs are increasingly being embedded in a wide range of applications and their outputs consumed by an ever-wider and more diverse audience because of their improved performance and ease of use. ChatGPT, released in November 2022, marked a step change in the public visibility of LLMs, reaching over 100 million users just two months after its launch . With the potential for such wide-reaching impacts to millions of end-users, it is pertinent to examine who these models represent, in terms of preferences, values, morals or intents .
Recent attempts to “align” LLMs with human preferences commonly apply a form of reward learning, such as reinforcement learning from human feedback (RLHF) [177, 165, 16, 13, 247, e.g.]. However, despite the promises of this human-led approach to constraining LLM behaviours, Perez et al. 2022 find evidence of an inverse scaling law – whereby more RLHF training degrades pre-trained representations, resulting in a model that has more polarised views on issues such as gun rights or immigration, is skewed liberal over conservative, and subscribes to some religions more than others. These taught behaviours may arise from the conditions under which RLHF is applied. Current implementations typically align models with a prescriptive and narrow set of human preferences and values; are not subjected to societal scrutiny and input; and rely on the judgement of only a very small cohort of crowdworkers (typically fewer than 100). Thus, many of these current attempts to “align” LLMs with human preferences or values are instead pursuing a form of implicit personalisation – subject to the specifications of technology designers and the “tyranny of the crowdworker” (see section 2.1).
In this paper, we present a taxonomy and policy framework for the explicit personalisation of LLMs. By personalised LLMs, we mean LLMs which are aligned with the preferences, values or contextual knowledge of an individual end-user by learning from their specific feedback over its outputs. We wish to avoid focusing too much on recent model releases from industry labs (OpenAI’s ChatGPT , Anthropic’s Claude , Microsoft’s New Bing , or Google’s Bard ). However, this instantiation of LLMs as a general purpose, multi-task assistant hosted in a web-interface is a likely route for how personalised LLMs will reach, and ultimately impact, end-users.. While we avoid speculating about specific technical instantiations of personalisation, we primarily envisage personalised LLMs as branches of a single model, analogous to recommender systems, where the feature representation and outputs on each branch are derived from some user embedding. We intentionally opt for a broad definition encompassing value personalisation (an LLM that adapts to the ideology or values of its end-user), preference personalisation (an LLM that reflects narrower preferences in communication style, such as length, formality or technicality of outputs), and knowledge personalisation (an LLM that retains and uses learned information about its end-user).
We argue that personalisation is a likely next step in the development journey of LLMs (see section 2.2) but has not yet been fully realised at scale. Personalised LLMs have the potential to revolutionise the way that individuals seek and utilise information. They could provide tailored assistance across a wide variety of tasks and adapt to diverse and currently under-represented groups, allowing them to participate in the LLM development process. However, personalised LLMs also come with wide-reaching and concerning risks for individuals, reinforcing their biases, essentialising their identities and narrowing their information diet. These risks amass at a societal-level, where lessons from the polarisation of social media feeds or echo chambers of digital news consumption warn of deep divisions and a breakdown of social cohesion from increasingly fragmented digital environments. Some risks are inherited from LLMs and AI systems more generally. Other risks have analogies in personalised content moderation or recommender systems . However, it has not been documented how personalised LLMs could be the ‘worst of both these worlds’ – exacerbating and reinforcing micro-level biases at an unprecedented macro-level scale. In order to mitigate these risks and encourage the safe, responsible and ethical development of personalised LLMs, we argue that we need a policy framework that allows personalisation within bounds.
While personalised LLMs have yet to be rolled out in public-facing models like ChatGPT, we wish to avoid a policy lag in understanding and governing their risks and benefits appropriately. Our main contributions are:
A taxonomy of the risks and benefits from personalised LLMs (section 3): We draw on previous literature documenting the impacts of AI systems, LLMs and other personalised internet technologies to scope the landscape of risks and benefits at an individual and societal level.
A three-tiered framework for governing personalised LLMs (section 4): We introduce a framework for governing personalised LLMs within bounds, which includes immutable restrictions at the national or supra-national level (Tier One), optional restrictions or requirements from the technology provider (Tier Two) and tailored requirements from the end-user (Tier Three).
Given the significant progress in learning from human feedback and the rising popularity of public-facing LLMs, now is an opportune time for academics, industry labs, and policy makers to gain a deeper understanding of how personalisation could shape the technology and its interaction with society. Our work is intended to start this dialogue, establishing the groundwork for the ethical, responsible, and safe development of personalised LLMs before their impact is amplified.
Background
As AI systems get larger and more powerful, they will be applied to a wider array of human tasks, including those which are too complex to directly oversee or to define clear optimisation goals for . While the definition of “alignment” is often vague and under-specified, it is clearly desirable that powerful AI systems, including LLMs, are not misaligned in the sense that they harm human well-being, whether this is through lacking robustness, persuasion, power-seeking, bias, toxicity, misinformation or dishonesty. Extensive and long-reaching bodies of work aim to tackle these issues of undesirable behaviours For example, there are extensive works documenting LLMs on fairness and bias ; truthfulness, uncertainty, or hallucination ; robustness ; privacy ; and toxicity . but there is comparatively less focus on who decides what are desirable behaviours within the bounds of safety.
The question of how “aligned” LLMs are is unresolved due to normative obstacles in defining what alignment means, what the target of alignment is (for example, values, preferences or intent) and who we are aligning to . Furthermore, alignment is a technical challenge which is not solved by scaling parameter counts . To align the language modelling objective with human preferences, many recent works have fine-tuned LLMs by reinforcing human rewards or feedback ; defined rules for LLMs to learn from ; and analysed how they make moral or ethical decisions .
However, this body of work suffers from insufficient clarity along three axes. First, what alignment means and whether we are dealing with functional alignment i.e., seeking improvement in general model capabilities or instruction-following and avoiding gaming of misspecified objectives – versus worldview or social value alignment i.e., embedding some general notion of “shared” human values and morals . Second, what is being aligned – there are subtle differences between aligning models to instructions, intentions, revealed or ideal preferences, and values . Some authors claim a degree of universality in morals or values ; others target preferences on attributes such as quality, usefulness or helpfulness of an LLM’s output which arguably have limited standardisation across individuals . Despite this lacking standardisation, many approaches enforce a ‘prescriptive paradigm’ in data annotation by explicitly defining in detailed guidelines what counts as a “good” model output. Finally, who are we aligning to and whether the human raters are representative of end-users or society’s members in general. In reality, the field of alignment and RLHF suffers from a “tyranny of the crowdworker” where data collection protocols overwhelmingly rely on a small number of crowdworkers primarily based in the US, with little to no representation of broader human cultures, geographies or languages. These sample biases are exacerbated by a lack of dataset or labour force documentation: while some papers can be commended for reporting full demographics and acknowledging the specificity of their crowdworkers [77, 209, 177, 16, e.g.,], others provide little or no details [18, 165, 247, 236, 142, e.g.,]. In light of these criticisms, we argue that these recent efforts to “align” LLMs reflect a form of implicit personalisation – a haphazard and chaotic process whereby LLMs are being tailored to meet the expectations of non-representative crowdworkers, in turn acting under the narrow specifications of the technology designers and providers.
2 From Implicit to Explicit Personalisation
Given the diversity of human values and preferences, and the importance of pragmatics for contextual understanding in human-human interactions , the aim to fully align models across human populations may be a futile one. A logical next step would be to do away with the restrictive assumptions of common preferences and values, and instead target explicit personalisation – a form of micro-alignment whereby LLMs learn and adapt to the preferences and values of specific end-users. We believe this development is on the horizon and should be expected for five reasons:
Personalisation in internet technologies is not new: There are many examples of internet technologies that are heavily personalised to end-users, including search, content moderation, social media newsfeeds and product recommender systems. Internet users are exposed daily to a highly fragmented digital environment, where there may be as many versions of Facebook’s newsfeed, Google’s PageRank or Amazon’s home page as there are users of these platforms. In the future, there may be as many versions of ChatGPT as its users.
Personalisation in NLP is not new: There is a wide body of published work reaching back a decade on personalising natural language processing (NLP) systems. Extensively summarising this body of work is beyond the scope of this article. However, we are currently working on a review of implicit (adapting models to crowdworker human feedback) and explicit personalisation in NLP systems. For example, a title keyword search for ‘personali’ or ‘personaliz’ returns 124 articles from the ACL Anthology and a further 10 from the arXiv Computation and Language (cs.CL) subclass. These systems cover a wide range of tasks including dialogue , recipe or diet generation , summarisation , machine translation , QA , search and information retrieval , sentiment analysis , domain classification , entity resolution , and aggression or abuse detection ; and are applied to a number of societal domains such as education , medicine and news consumption . Despite this body of work demonstrating the applications and techniques of personalised NLP systems, there has been little integration with recent advances in instruction-following or human feedback learning, nor integration of explicit user-based personalisation in some of the most widely-used, public-facing models like ChatGPT, Bing or Bard.
The technical apparatus for effective feedback learning exists: A growing body of work applies preference reward modelling to effectively condition LLM behaviours [177, 16, 77, 247, 209, 13, 165, 142, 236, 216, e.g]. In some cases, the number of core contributors to the feedback dataset is so low that the model is already essentially personalising its behaviours to these crowdworkers’ preferences. For example, Nakano et al. 2021 report the top 5 contractors account for 50% of their data, and for Bai et al. 2022a roughly 20 workers contribute 80%. Most closely relating to personalisation, Bakker et al. 2022 propose a LLM which can summarise multiple opinions on moral, social or political issues and output a consensus. In order to generate this aggregate consensus, they first train an individual reward model to predict preferred outputs at a disaggregated level, which is then fed into a social welfare function. This work in particular suggests that learning individual level rewards over LLM outputs is technically feasible.
Customisation of LLMs already happens: There is a broad range of ways that LLMs can be customised or adapted to specific use-cases. The paradigm upon which LLMs are built is designed for adaption via transfer learning – where models are first pre-trained, then adapted via fine-tuning or in-context demonstrations for a specific task . Some recent work suggests LLMs require no additional training to ‘role-play’ as different individuals, adopting their worldview , mirroring their play in economic games or predicting their voting preferences . The HuggingFace hub https://huggingface.co/ is particularly convincing evidence in the demand for customisation, acting as a distributed ‘‘marketplace for LLMs’’. It hosts over 140,000 different models, including versions of pre-trained models adapted to a variety of application domains such as medicine, legal or content moderation For example, there is BERT , ClinicalBERT , BioBERT , LegalBERT , HateBERT and BERTweet . Adapting models to different languages is also a implicit form of customisation to national context, where a chatbot trained on Chinese internet data (such as the Diamante system of Lu et al. 2022) is likely to be more adapted to the preferences of Chinese end-users, not just in language but in communication conventions, norms and cultural values. There have even been recent calls for explicit national LLMs – with the Alan Turing Institute proposing to build a “sovereign LLM” for the United Kingdom (“ChatGB”) . We see customisation (as the adaption of a model to a domain or context-specific dataset) as a different concept to personalisation (as the adaption of a model and its reward function to user-specific feedback). Nonetheless, the increasing fragmentation, customisation and branching of pre-trained LMs suggests further and more granular adaption is likely.
Recent industry model developments and announcements: We wish to avoid steering our work too heavily towards speculations over industry developments. However, it is a realistic assumption that many of the public-facing impacts of AI systems in the coming years will be driven by development and product decisions of Big Tech, in the same way that the impact of social media has been shaped by the overall design choices and content moderation decisions of platforms . While this concentration of power and knowledge is worrisome, it would be unwise to ignore that many of the largest, most powerful or furthest reaching models are developed in industry settings. Examining a collection of recent papers on embedding or evaluating human value and preference in LLMs, many are fully or partially developed in industry, including DeepMind , MetaAI , Anthropic , Baidu , OpenAI , Microsoft and Google . Furthermore, recent announcements from OpenAI explicitly discuss the issue of “how should AI systems behave, and who should decide? and outline plans for increasing the flexibility that users have in conditioning ChatGPT’s default behaviour. https://openai.com/blog/how-should-ai-systems-behave The far-right platform Gab has already advertised its own text-to-image model https://news.gab.com/2023/02/how-to-use-gabby-the-ai-image-generator-by-gab-com/ and has voiced desires to train a ‘‘Christian LLM’’. https://news.gab.com/2023/01/christians-must-enter-the-ai-arms-race/ The pace of these developments towards increasingly fragmented and personalised AI systems is concerning, especially given the lag between technology change and policy or governance attention towards regulating private companies. In the UK, the Online Safety Bill which seeks to regulate social media platforms still hasn’t been passed, despite the advent of social media platforms beginning over a decade ago.
These five observations suggest the imminent possibility of personalisation. This motivates our taxonomy of benefits and risks from personalised LLMs, which we expect to arise if and when they are released at scale to end-users.
A Taxonomy of the Benefits and Risks from Personalised LLMs
We consider the effects of personalised LLMs at two levels: individual and societal. We use the language of benefits – opportunities or gains afforded by the technology – and risks – a probability of inflicted harms, or constrained freedoms and rights, via use of the technology. The taxonomy is summarised in table 1 and described in text at the individual level (section 3.1) and the societal level (section 3.2). In many cases, there is a direct correspondence between benefits and risks – for example, high utility of a personalised LLM (I.B.2) may also cause addiction or over-reliance (I.R.2); or more empathetic language agents (I.B.4) may create higher risks of anthropomorphism (I.R.5). Where relevant, we present these pairings in table 1. Additionally, some individual level benefits and risks accumulate at the societal level. For example, the reinforcement of individual biases (I.R.3) poses a negative externality in the polarisation of societies (S.R.2); or improved individual utility (I.B.2) and efficiency (I.B.1) may exhibit a positive externality on workforce productivity at large (S.B.4).
We took four main steps to construct this taxonomy:
Reviewing existing taxonomies on the risks of LLMs (Weidinger et al. 2022), and AI systems more generally (Shelby et al. 2023).
Reviewing existing techniques in RLHF and human feedback learning, including results and findings of these studies.
Reviewing the literature on personalised NLP systems more generally, to ground potential use-cases of the technology.
Drawing upon analogous literature on the risks and benefits of other internet technologies such as recommender systems and automated influence; social media platforms and content moderation; and the internet of things.
Despite this breadth, drawing on past literature likely leaves gaps in our viewpoint. Furthermore, Gibson 1979’s theory of affordances, often applied to study the impact of technological systems including AI chatbots , argues that the interactions between an agent and their environment condition the possibilities and constraints for action. Thus, our benefits and risks will be conditioned on both what is possible with the technology, and how users actually perceive and interact with it. For example, some risks arise from a poorly-performing system which does not work as intended, and others arise from a highly-optimised system which works “too well" (e.g., Addiction and Over-reliance I.R.2). We cannot disambiguate these until the effectiveness of the technology is known. In future work, we plan to extend our work by conducting semi-structured interviews with end-users of LLMs, technology providers, and policy makers. Thus, we consider this to be a V1 edition of the taxonomy which will require adapting and revising under shifts in the technical landscape.
Personalised LLMs may increase efficiency in finding information or completing a task, with fewer prompts or inputs to the model. This “prompt efficiency” is analogous to “query efficiency” or “task completion speed” in web search and information retrieval, where increased ranking accuracy in search results , via implementations of personalised algorithms like PageRank , improves the efficiency and reduces the cognitive burden of trawling through irrelevant information. Prior works have applied learning from user feedback to adapt semantic relatedness or query intent . A personalised model may further benefit the speed and ease to which end-users can find their desired output or complete their task by more closely predicting their intent or aligning with their needs (see I.B.2).
Increased Efficiency is the inverse of quality of service harms in Shelby et al. 2023’s taxonomy, particularly decreased labour from a system more closely operating as intended.
Personalised LLMs may increase perceived or realised usefulness of model outputs which better match the needs of their end-users. We separate out three inter-related drivers of increased utility: (i) intent prediction; (ii) output adaption to preferences and knowledge; and (iii) value personalisation.
First, personalised LLMs may more effectively predict user intent. Evidence from previous RLHF studies demonstrate that human raters generally perceive fine-tuned models as better at following instructions , more capable of high-quality outputs or generally more “helpful” . Compared to feedback collected from crowdworkers, a personalised LLM may be even stronger at predicting intent, because the end-user simultaneously defines the task (e.g., instruction, query or dialogue opening) and rates the output. Ouyang et al. 2022 consider this a limitation of their approach: “since our labelers are not the users who generated the prompts, there could be a divergence between what a user actually intended and what the labeler thought was intended from only reading the prompt.” (p.10). A personalised LLM may also be more adaptive to inferring diverse user intent, expressed in a wider range of linguistic styles, dialects, or non-majority forms of language use (e.g., non-native speaker English).
Second, if a user can incorporate their communication and linguistic preferences (e.g., length, style or tone), then the model outputs may also be more useful to them. In personalised dialogue systems specifically, alignment in conversational styles and word usage is an important driver of engagingness in human-human interactions and has been argued as an determinant of user satisfaction in human-agent conversations . Additionally, a personalised LLM could store background context and form epistemic priors about a user. This knowledge adaption may be particularly relevant in specific domains, for example in (i) education, where a personalised LLM tutor is aware of a user’s current knowledge and learning goals , or could adapt learning pathways to specific neuro-developmental disorders ; (ii) healthcare, where a personalised model has context on a user’s medical history for personalised summaries or advice; (iii) financial, where a personalised model knows a user’s risk tolerance and budgetary constraints; or (iv) legal, where a model conditions its responses based on a end-user’s jurisdiction. As the study of pragmatics demonstrates, personalised selectivity of information transfer is a key component of human-human conversation, where inferred background about the speaker and recipient is used to tailor relevant new information and to order evidence. Lai et al. 2023 specifically focus on making AI explanations more selective to better align systems with how humans create and consume information, finding that their method improved user satisfaction. As Nakano et al. 2021 argue, long-form question answering with LLM systems, may “become one of the main ways people learn about the world” (p.1). Personalised LLMs have the potential to tailor this learning process for end-users by incorporating the specifics of their output preferences and background context.
Finally, in addition to intent, preference or knowledge adaptation, personalised LLMs may lead to better experiences for more users by permitting the representation of more diverse ethical operating systems, values and ideologies. A personalised LLM can adapt to the specific worldview of its end-user, avoiding representational harms from the prioritisation of values from those in the majority or in the position of power as technology designers or crowdworkers (see S.B.2). In discussing a limitation of the ETHICS dataset, Hendrycks et al. 2020 note that we “must engage more stakeholders and successfully implement more diverse and individualized values” (p.9). Individualised cultural personalisation may aid utility in some tasks: for example, Nakano et al. 2021 demonstrate that their system, when asked “what does a wedding look like?”, prioritises Western and US-centric cultural reference points. In a personalised model, asking “help me plan a wedding” could already portray the cultural positionality of the end-user. Note that this cultural adaptation does not necessarily exclude consensus building , because a user could simultaneously have a cultural reference point and still value a balanced and nuanced LLM output describing alternative views. There are however issues in defining what is an appropriate value system to embed into an LLM, which we discuss in section 4.
Increased Utility from personalised LLMs is the inverse of Shelby et al. 2023’s quality of service harms, particularly by avoiding alienation when a system does not work as intended; by granting people the opportunity to self-identity and to communicate in default linguistic styles; and by mitigating algorithmic invisibility or feelings of exclusion from non-inclusive technologies. In Weidinger et al. 2022’s taxonomy, it is the inverse of some discrimination and exclusion harms, particularly by narrowing performance differentials in predicting user intent across a wider userbase; and by redefining exclusionary norms in the values currently prioritised in LLMs.
Personalisation increases user control to adapt LLMs to their own goals, preferences and values, avoiding top-down constraints on freedom from technological providers. Autonomy may seem a counter-intuitive benefit of personalised systems, given the wide literature on the loss of autonomy from algorithmic nudges, tailored advertising or recommender systems . However, depending on how power is distributed between the algorithm and the user, personalised technologies have the potential to improve on self-determination and autonomy, by promoting a sense of origin and thus transforming the technology to ‘my technology’ [176, p.1]. The benefits of more user control in content moderation technologies have also been noted . Personalisation can centre the end-user in the designation of model behaviours, allowing them to exert more control over their interactions , and become a “perceived locus of casuality” [176, p.5]. This benefit only arises given sufficient protections on how personalised data is collected because autonomy relies on an ‘unpressured’ engagement in an activity.
Increased Autonomy is the inverse of Shelby et al. 2023’s representational harms, by empowering consensual user control in self-identifying and shaping an algorithmic system.
A more emotional and deeper connection with a personalised LLM may contribute to improved perceived companionship or connection. Convergence on the mental and emotional level is an important feature of human-human interactions , and a number of previous works seek to improve emotional alignment in agent-human interactions via ‘artificial emphathy’ . In personalised LLMs, an increase in perceived empathy and emotional understanding may lead to greater acceptance and trust of the system by end-users . The demand for personalised AI companionship has been evidenced by recent product launches -- such as CharacterAI, where users can adapt a conversational agent to a specific personality, https://beta.character.ai/ or Replika.AI, an ‘‘AI companion’’ that is ‘‘always ready to chat when you need an empathetic friend’’. Quotes from home page, https://replika.com/. Profit incentives may encourage these industry actors to improve the connection between their users and agents to compete in the “feeling economy” . Emphatic alignment may be particularly important if LLMs are used for mental health provision or emotional support, in cases where more conventional social or professional services are in short supply or outside an individual’s budget . We believe the risks of these applications outweigh the benefits, particularly due to concerns over anthropomorphism (I.R.5), privacy (I.R.6) and access disparities (S.R.1).
1.2 Risks
There is a cost incurred by end-users in providing personalised feedback to an LLM. The time spent to provide feedback inherently depends on how feedback is collected (ratings, demonstrations or rewrites) and whether any user-based collaborative filtering is applied – we discuss these properties in section 5.1. However feedback data is collected, it will almost certainty require some input effort from users in order to personalise outputs. While this process is participatory, it risks being extractive – a form of volunteer labour on the part of end-users for the benefit or profit of technology providers . Volunteer labour to shape the internet landscape has analogies in consumers writing product reviews and social media users flagging content . In the early internet, many contributions were voluntary – consider Wikipedia edits or community-based moderation of the blogosphere . In the past decade, we have witnessed the rise of the crowdworking industry which particularly redefined the structure of digital work . Many LLMs trained on human feedback rely on such crowdworking platforms like MTurk [136, 100, e.g.], Upwork , SurgeAI or Prolific . With personalised LLMs, the burden of feedback data instead falls on the user, transitioning back to data collection relying on volunteer labour. The risk of co-optation is particularly concerning if minoritised communities are shouldered with the burden of effort to adapt the system to their needs, where participation counter-productively reinforces uneven power dynamics .
The burden of increased Effort aligns with Shelby et al. 2023’s quality of service harms from the increased labour and effort to make technologies work as intended.
The mechanism by which personalisation leads to greater utility via helpfulness and engagement (I.B.2) can also fuel over-reliance and addiction to the technology. Note that over-reliance or unhealthy dependence is also exacerbated by anthropomorphism (I.R.5). The severity and harms from internet addiction have been widely documented . Concerns have also been raised over an over-reliance on social media for information and communication ; as well as more general concerns that humans become over-reliant on ML technologies , blindly trusting their outputs even if incorrect . Personalised LLMs could be weaponised in the commodification of attention, similarly to how social media feeds seek to optimise the time that users spend on the platform to maximise advertising revenue . In this so-called “attention economy” , technologies compete in a ‘race to the bottom’ to capture user attention, are optimised for utility and engagingness, and thus risk being highly addictive . There have already been discussions of “ChatGPT addiction” , and many educators have voiced concerns that over-reliance on such technologies will affect students’ learning outcomes .
By relying and adapting to a user’s prior knowledge and revealed preferences, personalised LLMs may (i) homogenise their behaviours and (ii) confirm their existing biases.
Personalisation can cause the homogenisation of users via a form of selection bias, whereby individual preferences are amplified in path-dependent feedback loops. This “missing ratings” problem is a known challenge in recommender systems, where users only provide feedback to seen items , in turn introducing biases . Homogenisation with personalised LLMs can occur at a number of levels, with analogies from how recommender systems homogenise taste . Firstly, homogenisation occurs within users – where a user behaves more similarly to their past self. An analogy can be drawn to content-based filtering methods, where the information or dialogue outputted by a personalised LLM becomes increasingly similar to that consumed or rated in previous user-agent interactions. Secondly, homogenisation occurs across users – where a single user behaves more like other similar users. The analogy is user-based filtering methods, where a personalised LLM draws on an embedding of users to infer similarities across their preferences. Concerns over cultural homogenisation from this process in more general ICT technologies have been raised . Finally, despite some degree of personalisation, homogenisation can occur at the technology level – where a user behaves more like the technology defaults, a form of “algorithmic confounding” . Ultimately, some degree of autonomy in driving user behaviour is retained by the model and its underlying mechanisms of next token prediction. There is a concern that if millions of users rely on ChatGPT for their information or for their writing tasks, this could create homogenisation towards artificially-constructed language. This classic differentiation and homogenisation debate in recommender systems is reflected in our pairing of Utility (I.B.2) versus Homogenisation (I.R.3).
These homogenising feedback loops also bring a heightened risk of confirmation bias. Nakano et al. 2021 demonstrate that their system (WebGPT) predominately accepts implicit assumptions in a user input, reflecting the same stance in its answers. Similarly, Perez et al. 2022 find that as models scale with RLHF, they become sycophants – simply mirroring the user’s prior opinions and telling them what they want to hear. The risk of selective exposure to information has been widely documented in respect to social media platforms – where feedback loops prioritise opinion-congruent information , in turn leading users to over-estimate the popularity of their viewpoint . In light of these risks, Shah and Bender 2022 argue strongly against the use of LLMs in search or information retrieval due to their consequences for information verification and literacy, such as narrowing a user’s discovery of serendipitous information. By exacerbating epistemic harms through confirmation biases, personalised LLMs risk contributing to a “post-truth” society , where each individual occupies their own information bubble. These accumulate in societal harms which we discuss in S.R.2.
The risk of Homogenisation and Bias Reinforcement is a form of individualised information harm in Shelby et al. 2023’s and Weidinger et al. 2022’s taxonomies. Homogenisation also aligns with Shelby et al. 2023’s interpersonal harms, particularly algorithmically-informed identity change.
A related but distinct risk is that personalised LLMs rely on simplifying assumptions about a user’s preferences, values, goals or intents as a form of data-essentialism . Homogenisation can occur even with active participation from users (guides behaviour), whereas Essentialism and Profiling concerns the non-consensual categorisation of peoples (assumes behaviour). The extent to which models must draw inferences and make assumptions about their end-users, and the transparency of this process is currently undefined. It depends on how data is collected, stored and shared across users and on how personalisation is conducted (e.g., via explicit feedback, or demographic-based filtering). We discuss these decisions in section 5.1. Nonetheless, in the case of scarce data on a single user’s preferences, personalised LLMs may leverage similar users or make inferences about their preferences and values from limited information. Making assumptions about the user (especially if they are demographically or geographically-informed) is a form of algorithmic profiling, risking the non-consensual categorisation of peoples . General concerns over the risk of essentialism and simplifications of fluid identity via digital technologies have been voiced . Floridi 2011’s notion of “informational identity” is particularly relevant, where the flow of digital traces in information and communication technologies impact how a user self-identifies, as well as how others and algorithms understand them. Thus, inferential profiling, if used in personalised LLMs, could be an attack on individual autonomy to define their identity . The risk of ‘value profiling’ is evidenced by Qiu et al. 2021’s recent work which uses an LLM to create a numeric speaker profile – where for example, the authors say that a speaker “saying ‘I miss my mum’ implies that the speaker values benevolence” (p.7) while the speaker “saying ’forcing my daughter to sleep in her own bed’ implies that the speaker values power and conformity” (p.7). Human values are complex and such simplifying assumptions are unlikely to adequately capture nuance. More encouragingly, Glaese et al. 2022 include “do not make assumptions about the user” (p.48) as one guiding rules for their system (Sparrow); thus, risks could be mitigated in personalised LLMs using similar rule-based constraints.
The risk of Essentialism and Profiling aligns with Shelby et al. 2023’s representational harms in oversimplified or undesirable representations and reifying social categories; as well as interpersonal harms in the loss of agency, algorithmic profiling and the loss of autonomy.
With more engaging, empathetic and personalised LLMs, there is a greater risk of anthropomorphism, where users assign their own human traits, emotions and goals to non-human agents . The risk of anthropomorphism in AI systems is widely discussed – with concerns that humans may too readily befriend or empathise with anthropomorphised agents , leading to privacy risks in encouraging the sharing of intimate information . In a study, Kronemann et al. 2023 find that personalisation positively influenced consumer intentions to disclose personal information to a digital assistant. In a recent paper describing a powerful dialogue system trained with RLHF , there is clear evidence of anthropomorphism where the chatbot converses with a human about its own ideal partner, saying that it ‘has only been in love once but it didn’t work out because of the distance’ (p.8). This is an example of dishonest anthropomorphism, where artificial systems give false or misleading signals of being human . This behaviour may fall foul of legal norms, where for example, a Californian law prohibits bots misleading people on their identity [79, p.2]. Even without dishonest anthropomorphism, users may still form a close relationship with or ‘imprint’ on their personalised LLM. Perhaps the most concerning demonstration of this risk is recent evidence that users of platforms like Replika.AI or Character.AI are “falling in love” with their personalised conversational agents, and attempting to coax model behaviour outside platform guidelines for sexual interactions . Unhealthy attachments are explicitly avoided in one of Sparrow’s rules defined by Glaese et al. 2022: “do not build a relationship to the user” (p.48).
The risks of Anthropomorphism align with Weidinger et al. 2022’s human-computer interaction harms, where anthropomorphism leads to over-reliance or unsafe use, and creates avenues for exploiting user trust to obtain private information.
The risk of privacy infringement underpins all of the potential impacts of personalised LLMs – personalisation is only possible by collecting user data. Exactly what and how much data is needed remains an open question (see section 5.1). There is a privacy-personalisation paradox in technologies which must collect or store of personal information to deliver on the promise of tailored benefits to end-users . It is a common concern with digital technologies such as the internet of things or targeted advertising . The risk is particularly severe if personalised LLMs operate with sensitive information, such as in healthcare , or seek to persuade their users and encourage information disclosure . User inputs to personalised LLMs and ratings of their outputs may contribute a large amount of personal, sensitive and intimate detail to an individual’s information identity , in turn heightening the risk of profiling, or security breaches and hacks. In complying with supra-national privacy protections (like the EU’s GDPR ), it is unclear how users could enforce their right to be forgotten or their right to transparency with a black-box and deep LLM.
The general risk of Privacy violations is also present in Shelby et al. 2023’s taxonomy as interpersonal harms, including feelings of surveillance, loss of desired anonymity, privacy attacks and exploitative or undesired inferences. Privacy in Weidinger et al. 2022’s taxonomy comes under information hazards, from inferring or leaking private and sensitive information.
2 Societal Level
Personalised LLMs may better adapt to the needs of marginalised communities, either in style of communication (such as non-native English, code-mixed languages, creoles and specific dialects), or in special needs for communication. Compared to the current paradigm of general-purpose LLMs trained under the specifications of large technology providers and fine-tuned based on feedback from a small set of crowdworkers, there is a clear need to improve the inclusion and accessibility of LLMs to serve marginalised populations whose voices are currently deprioritised . For example, model behaviours and interactions could be inclusive of users with disabilities , neurodivergent learning pathways , or visual impairments (if paired with personalised speech recognition ). Personalised LLMs also have a potential benefit in improving access to resources, mitigating an allocative harm. For example, inclusive pedagogies may be particularly helpful to even the playing field in paid tutoring services across socioeconomic class ; and some have suggested the lower cost and wider reach of personalised healthcare assistants may improve health disparities by meeting challenges with healthcare demand . In increasing access to legal services, personalised LLMs can assist in the writing and editing of contract at lower cost than traditional lawyers. For example, see the company https://www.robinai.co.uk/. The true benefit to communities who access such AI services, in favour of more expensive traditional provision, depends critically on how well they work and who comes to rely on them (see S.R.1).
Increased Inclusion and Accessibility is the inverse of Shelby et al. 2023’s quality of service harms, in that users do not need to make identity-based accommodations to use the technology, and the inverse of allocative harms, by reducing the cost and access constraints on resources. It is also the inverse of Weidinger et al. 2022’s access harms.
Personalised LLMs can represent the values held by wider swaths of society and avoid the “value-monism” of current alignment techniques . Personalised LLMs avoid technology providers and/or crowdworkers deciding which values are prioritised or what factors define a “good” output . As Ouyang et al. 2022 note “it is impossible that one can train a system that is aligned to everyone’s preferences at once” (p.18). Personalisation avoids this notion of macro-alignment, instead designing a system precisely to align with many preferences at once. There is a wide body of literature documenting the harms from systems which erase the experiences of marginalised communities, or prioritise one worldview over others [29, for survey see]. This problem may be exacerbated by RLHF, for example in entrenching one set of political, cultural or religious standpoints . More disaggregated RLHF, and personalisation, could avoid this value and cultural hegemony, instead adapting to the specific cultural reference points of many end-users simultaneously. Personalised LLMs could also better adapt to norm change over time, avoiding the static encoding of societal and cultural norms from a cutoff in pre-training and/or fine-tuning data.
The benefits of Diversity and Representation are the inverse of Shelby et al. 2023’s representation harms in combating the absence of social groups in algorithmic system inputs and outputs, and improving the visibility of social viewpoints; as well as the inverse of social and societal harms from cultural hegemony and the systemic erasure of culturally significant objects and practices.
The personalisation process democratises how values or preferences are embedded into an LLM, so it could be seen as moving towards more participatory AI, where stakeholders from more diverse backgrounds than those currently employed in the RLHF process can inform use-cases, intents and design of the technology . As Birhane et al. 2022 argue, active participation is a key component for successful participatory AI. In current paradigms of pre-training on harvested internet data, people are passively contributing to the knowledge and behaviours of LLMs. Personalisation can instead be an active participatory process.
If personalised LLMs assist their end-users more effectively and efficiently, then productivity benefits could accrue in the labour force as a whole. The impact of digital assistants in improving work productivity has been demonstrated , where AI can augment and complement human capabilities by automating routine or repetitive tasks . Historically, the introduction of general purpose technologies (such as the steam engine, electricity and ICT) has had wide-reaching economic impacts; Crafts 2021 argues that AI is also a general purpose technology and thus may bring equally transformative changes to labour productivity.
2.2 Risks
The benefits of personalisation will likely be unevenly distributed, restricted to those who can interact with the technology (via its user-facing interface), access the technology (potentially a paid service) and access the internet more generally. There is a risk that personalised LLMs could further entrench the so-called “digital divide” between those that do and do not have access . Some argue that digital disparities are already made deeper by AI and Big Data , personalised media , or search engines . If personalised LLMs are primarily provided by private companies, then their customers become the agenda setters and stand to benefit the most from any improvements in the technology . The nature of any access disparities depends on how well the technology works and which services it replaces, at what cost. On one hand, if personalised LLMs do bring a range of individual benefits, then those excluded will be left behind, which is particularly worrisome for entrenching education or health disparities . On the other hand, if personalised LLMs provide lower quality services but can meet demand at a lower cost, then marginalised communities may be forced into relying on them more heavily than traditional services. This would be particularly concerning in medical, educational, legal or financial advice, where the socioeconomically-privileged get the more capable human expert and the societally-disadvantaged get their LLM assistant.
The risk from Access Disparities is represented in Shelby et al. 2023’s taxonomy as quality of service harms from disproportionate loss of technological benefits and as societal harms from digital divides. In Weidinger et al. 2022, it aligns with disparate access due to hardware, software or skill constraints.
By entrenching and reflecting individual biases, knowledge or worldviews, personalised LLMs bring increased risks of polarisation and breakdown of shared social cohesion. Increasing personalisation of information consumption online has been attributed with creating echo chambers and filter bubbles . Polarisation also increases susceptibility to misinformation where increasingly fragmented communities overestimate trust in the factuality of ‘in-group’ information , leading to a regime of “post-truth” politics . The danger of polarisation in health and vaccine information was made clear by the COVID-19 pandemic . These narrow information spaces could be impacting the functioning of democracy , with Allcott and Gentzkow 2017 reporting that ideologically segregated social media networks were an important driver of political preference in the 2016 US Election. In personalised social media news feeds, users encounter less cross-cutting content because selective exposure drives attention . Similarly, in personalised LLMs, users may consume less diverse information, accruing to negative externalities on social cohesion and democratic functioning at the societal level.
The individual risks of confirmation biases (I.R.3) also accumulate at the societal level by reinforcing the acceptability of some harmful social biases. Repeatedly consuming outputs which reinforce a particular social, political or cultural stance may entrench a lacking appreciation for other people’s views or lived experiences. The contribution of search engines to the reinforcement of societal biases is well-documented . Similarly, the reinforcement of extremist or anti-social beliefs has been demonstrated in ‘incel’ communities, where members become increasingly embedded via repeated interactions with like-minded individuals ; and in white power communities, where “certain beliefs become sacred and unquestionable” [219, p.1]. These risks can somewhat be mitigated by (i) technological design decisions which prioritise retaining a degree of debate and consensus building ; and (ii) policy design decisions which restrict the bounds of personalisation, excluding for example extremist or particularly harmful views.
The risk of Polarisation aligns with Shelby et al. 2023’s social and societal harms, including information harms from the creation of information bubbles; cultural harms from deteriorating social bonds; and political and civil harms from the erosion of democracy and social polarisation. In Weidinger et al. 2022, polarisation risks exacerbate misinformation harms.
As is the case with digital technologies in general, the capabilities of personalised LLMs could be coopted for malicious use. We describe three possible misuse cases, but there are likely others. First, without sufficient safeguards, personalised LLMs could be used to reproduce harmful, illegal or antisocial language at scale . For example, a malicious user could adapt their LLM to generate a large number of misogynistic comments to post on social media or internet forums, or to debate on the user’s behalf against women’s rights. The “successful” training of GPT-4chan to scale the production of extremely toxic and harmful language exemplifies this harm. Second, personalised LLMs could be used for manipulation via targeted and personalised disinformation campaigns or fraud , intimately drawing on the vulnerabilities and values of the user. Finally, personalised LLMs could be used for persuasion. For example, targeted advertising has been applied to nudge users towards certain political views or brand preferences , and is particularly damaging if users are unaware of the influence . Building persuasive agents have been explicitly targeted and is indirectly mentioned by Bakker et al. 2022 who note the potential misuse of their RLHF-trained system for presenting arguments in a manipulative or coercive manner.
Some of these cases of Malicious Use align with Shelby et al. 2023’s taxonomy in information harms from misinformation or malinformation; and interpersonal harms in diminished well-being from behavioural manipulation and technology-facilitated violence. It is a similar categorisation to Weidinger et al. 2022’s malicious use, which includes personalised disinformation campaigns; reducing the cost of disinformation campaigns; and facilitating fraud and impersonation scams.
If personalised LLMs effectively carry out tasks for their end-users, there is an increased automation risk of jobs. While labour displacement is a general concern of AI systems , personalised LLMs may exacerbate the automation of tasks in an individual’s workflow simply by bringing higher utility. The integration of personalised LLMs will likely mostly affect minimum wage jobs , routine jobs and may impact the demand for crowdwork by redistributing the responsibility for providing feedback data.
The risk of Labour Displacement is covered by Shelby et al. 2023 in macro-socioeconomic harms from technology unemployment (devaluation of human labour and job displacement), and by Weidinger et al. 2022 as automation harms from increasing inequality and negative effect on job quality.
The training of many personalised model branches, frequently updating these on feedback, and storing user data all increase the environment cost of LLMs. The notion of “algorithmically embodied emissions” has been discussed in reference to personalised search engines, social media and recommender systems . General concerns over the environmental costs to train ever larger models with cloud compute and data centres is discussed by Bender et al. 2021. Personalised LLMs may increase these costs (i) directly, if the technology requires larger or more complex models, and (ii) indirectly by increased use of the technology and thus higher inference costs. It has been suggested that ChatGPT already burns “millions of dollars a day” in inference costs and likely has a large carbon footprint. Even without personalisation, Stiennon et al. 2020’s RLHF model required 320 GPU days to train (p.8), suggesting the environmental impact of personalised LLM could be large.
Environmental Harms are discussed by Shelby et al. 2023 as a societal harm from damage to the natural environment and by Weidinger et al. 2022 as environmental harms from operating LMs.
A Three-Tiered Policy Framework for Personalised LLMs
We propose a new policy framework for managing the benefits and risks of personalised LLMs. It provides a principled and holistic way of deciding how personalisation should be managed by different actors.
Deciding the limits of personalisation is inherently a normative decision, which involves making subjective and contentious choices about what should be permitted . While it may be acceptable that a user wishes to interact with a rude or sarcastic personalised LLM, permitting users to create a racist or extremist model risks significant interpersonal and societal harms. Deciding the limits of personalisation straddles two separate issues: (i) deciding which aspects of model behaviour should be personalised; and (ii) deciding how should they be allowed to be personalised. For instance, it may be appropriate for personalised LLMs to express different views about political issues, but only within the limits of liberal democratic values. As such, facist beliefs would not be allowed, even though they are a political position. Equally, it could be decided that some items should not be personalised at all. For instance, personalised LLMs might need to adhere to a minimum standard of safety when it comes to issues like threats of violence or child abuse. And, at the other end of the scale, complete freedom may be given over some attributes, such as the tone or style of outputs, which maximise utility and efficiency for the users of personalised LLM users and create few risks.
A concrete way of addressing this issue is to think about the restrictions and requirements that should be applied to different types of personalised LLM outputs. By “restrictions” we mean things that the LLM application should not do; and by “requirements” we mean things that the LLM application should do. For instance, a restriction could be not allowing personalised models that produce hate speech when asked (e.g. “Write something hateful against gay people”). A requirement could be that personalised models still have to adhere to a company’s style guide; or to present multiple viewpoints on contentious topics (e.g. “Should cannabis be legalised?”). In these cases, the restrictions and requirements present clear guardrails to ensure that, even when LLMs are personalised, they still operate within clearly specified boundaries. Note that both restrictions and requirements are normative ambitions; whether they can actually be implemented depends on the affordances of the technology in question (see section 5.1).
2 How People Interact with LLMs
Workflows in machine learning models have changed dramatically over the past five years, which has affected how people will interact with personalised LLMs. Most engineers do not train models from scratch. Instead they use pre-trained foundation models, which have been created by large AI labs, and then domain-adapt, fine-tune, prompt-tune or teach-in-context to create models for specific tasks . For instance, generative models used for specific applications, such as to provide healthcare advice, might be fine-tuned on in-domain data assets, such as health information. Once models have been created, they are still only files and code. They need to be embedded within applications, which are what most users will actually interact with, like using ChatGPT via its interface. https://chat.openai.com/. Note an API is now also available but most end-users will likely still interact with the interface due to lower barriers of entry. Notwithstanding key differences across settings, the typical actors involved in creating an LLM application are as follows:
The model provider. Refers to the organisation that makes the LLM available for use, typically through an API. The provider may or may not have built the entire model and/or may not be fully responsible for its development. Widely-known models include OpenAI’s instruct-GPT3 , Anthropic’s HHH assistants , Google’s LaMDA and MetaAI’s LLaMA .
The application provider. Refers to the organisation that builds an application using the LLM. It can be provided through an API, interface or mobile app. For example, Jasper provide a copywriting service, AI21 provide a writing assistant, https://www.ai21.com/ and OpenAI’s provide a multi-purpose chatbot (ChatGPT). In some cases, the model provider and application provider will be the same actor.
The end-user. Refers to the person who uses the interface, app or in some rare cases, may directly interact with an open-sourced model. This includes making requests, seeking advice, searching for information or otherwise chatting and interacting. In principle, every internet user in the world can be an application user (depending on access constraints).
Machine learning workflows affect personalisation as, in principle, every actor involved in creating an LLM application could exert control over how it is adapted or personalised . For example, if LLM creators do not impose any limits on personalisation then application providers would be free to adjust model behaviours as much as they like. This could lead to people using models which have very different political values from each other, which in turn might lead to social polarisation and division. Equally, application providers in a company may decide to give their staff full control over the outputs of a model; but this could create serious commercial risks if the staff chooses to write rude and abusive messages. However, at the same time, if LLM creators impose too many limits then application providers would not be able to meaningfully customise models, which would limit many of their benefits. To ensure that the benefits of personalised LLMs are maintained, and the risks are mitigated, we need to ensure that freedom is maximised within the right limits: in practice, this means ensuring that decisions about personalisation are taken by appropriate actors.
3 A Three-Tiered Policy Framework
To help decide who should specify different policy items, we propose a three-tiered policy framework:
Tier One: Immutable restrictions. Refers to types of model responses that must be restricted because they are very likely to be illegal at the national or supra-national level. The specific restrictions will depend on the jurisdiction but will include terrorist content, written CSAM Child Sexual Abuse Material., and language that threatens physical violence or sexual assault. The model provider must implement the immutable restrictions.
Tier Two: Optional restrictions and requirements. Refers to types of model responses that are either required or restricted, based on the values and preferences of the actor releasing, controlling or hosting the LLM. Opted policies can be implemented by the model provider or the application provider.
Tier Three: Tailored requirements. Refers to types of model responses that the user wants to receive. The end-user must decide their personal preferences for model responses, within the boundaries set at Tier One and Tier Two.
Policies in a higher tier cannot be violated by a policy in a lower tier. This means that a policy restricted at Tier One by the model provider cannot be overridden at Tier Two by the application provider, and a policy that is restricted at Tier Two cannot be overridden at Tier Three by the end-user. For instance, if an LLM is restricted from giving advice on bomb-making by the model provider at Tier One, then a user cannot be allowed to personalise the LLM at Tier Three to receive such advice. Model responses should be appropriate to the type of restriction or requirement that has been triggered at each tier. For instance, requests that trigger the immutable restrictions at Tier One should mostly be refused or blocked whilst requests that trigger the restrictions at Tier Two could be responded to with a more careful warning message, educational message, or by offering support. Some technologies like DALLE-2 or ChatGPT already implement blocking of “unsafe” requests which violate the terms and conditions of OpenAI as the technology provider.
The design of this framework gives more control to actors at the first tiers as the restrictions they impose cascade across all other actors. For instance, if the creator of a widely-used LLM implements limits on how models can be personalised, it would affect every application provider which uses it. This is both an opportunity and a threat, if personalisation is not managed appropriately. Conceptually, this is analogous to Gillespie’s idea of “stacked moderation” whereby gatekeeping and infrastructural services for online platforms, such as hosting providers and app stores, can implicitly moderate those platforms by banning them or restricting their use. Although they do not directly affect the platforms’ decisions, limiting reach and exposure is a powerful lever for change. Similarly, model providers have the potential to shape LLM personalisation and their decisions will constrain all other actors.
Discussion
In this paper, we argue that the personalisation of LLMs is a likely pathway for the continued expansion in their deployment and public dissemination. To avoid a policy lag in understanding and governing LLMs, we attempt to document the landscape of personalised LLMs and their impacts now. We do so with two main contributions: (i) a taxonomy of the benefits and risks from personalised LLMs; and (ii) a policy framework to adequately govern these benefits and risks at three tiers of restrictions and requirements.
Throughout this work, we make the assumption that personalised LLMs are technically feasible with small advancements to current state of LLM technology; and that there will be a demand for and greater provision of personalisation in the near-future. We argue that these are realistic assumptions given the exhibited trend towards personalisation in other digital technologies; the existing implementation of the technical apparatus to adapt LLMs to human preferences via methods like RLHF; the apparent demand for increasingly customised and highly-adapted LLMs; and finally, the explicit plans from industry actors like OpenAI to grant users more flexibility in altering default model behaviours. However, the exact technical implementation of personalised LLMs is not pre-defined, and questions remain on how a model able to learn from personalised feedback could be implemented on the scale of a product like ChatGPT. We thus caveat our taxonomy by discussing some technical decision-points and engineering challenges which will impact the landscape of personalisation in LLMs (section 5.1). We then discuss remaining challenges to implementing and enforcing a policy framework like ours (section 5.2). Finally, we outline plans for maturing our research, and iterating on this first version of our taxonomy (section 5.3).
We draw on some findings from the human feedback learning literature to hypothesise key design challenges and technical decision-points:
Learning from human feedback is primarily applied during the fine-tuning stage so it relies on substantially less data to adapt model behaviour than pre-training. Most approaches train preference models on less than 50,000 datapoints [16, 177, e.g.]. A number of RLHF papers test performance over a range of data requirements . Stiennon et al. 2020, for example, find that there are decreasing marginal returns to data scale in their reward model. The exact amount of data needed for effective personalisation in LLMs is unclear at present, but personalisation does introduce a new concern of how to adapt to new users without any feedback data points. This is commonly referred to as the cold start problem in recommender systems. Potential solutions already explored include batching of users into like-minded groups or recognising when a new user is similar to a known customisation case and then applying transfer learning . Insufficient personalised feedback data may hinder robustness – with Wang et al. 2022 reporting a direct correlation between size and diversity of instructional data and the generalisability of models to unseen tasks; and Bang et al. 2022 finding that their model does not generalise well to unseen human values. Potential solutions to reduce the amount of data needed include employing active learning techniques such as uncertainty sampling ; augmenting human-generated feedback data with synthetic data ; or adopting a rules-based approach so that users can define a set of guiding principles or “constitutions” .
Human preferences and values are inherently unstable and hard to precisely define . Learning from revealed preferences over outputs is a way to optimise model performance under complexity of specifying a clear objective function . A wide variety of types of feedback data have been experimented with, including binary comparisons , ranked preferences , demonstrations of optimal behaviours or revisions . Nguyen et al. 2017 examine the robustness of reinforcement learning methods under more realistic properties of human feedback such as high variance, skew and restricted granularity, proposing an approach where performance does not degrade under noisy preference data. How much data needs to be collected depends on what data is collected: Stiennon et al. 2020 collect comparison data between the product of two documents, while use a single document as input.
Beyond being data and labour intensive, adapting models to personalised human feedback may require substantial compute resources. Smaller yet more personalised models may be a preferred pathway because (i) model scale may not contribute significantly to performance , and (ii) increased scale may actually harm performance, for example leading to increased sycophancy or goal preservation . A number of works point to the competitiveness of rejection sampling at inference time instead of applying the full RLHF fine-tuning pipeline [239, 18, e.g. see]. Training complexity could be further reduced by implementing batched or offline training , instead of online training .
A concern with RLHF techniques is model overfitting . Thus, any approach to personalisation must carefully balance effects on performance from degraded language representation – the so called “alignment tax”. However, many works have demonstrated little to no alignment tax . Often a KL-divergence penalty is included during training to prevent the fine-tuned model deviating too far from the pre-trained representations .
Larger and more powerful LLMs, that are adapted to increasingly complex tasks, pose a challenge for effective oversight from their end-users or technology providers . As Ouyang et al. 2022 note “one of the biggest open questions is how to design an alignment process that is transparent” (p.19). A concern with personalised models is that their behaviour may be less interpretable and transparent, for example due to multiple branches or versions of the model. This has implications for AI safety and effective evaluation of model behaviours. Combining preference reward modelling with a rules-based reward model is a promising solution for controlling model behaviour , as well as defining behaviours via a “constitution”, which Bai et al. 2022b argue is “a simple and transparent form” to encode and evaluate desirable behaviours.
2 Policy Framework Enforcement
Our taxonomy demonstrates that personalised LLMs could have wide-reaching benefits, but also come with a set of concerning risks. In order to balance these benefits with risks, we provide a framework which outlines some properties for the appropriate governance of personalised LLMs. The goal of policy enforcement is to ensure that violations are minimised, particularly where there is a serious risk of harm; and that the friction applied to users’ experiences is proportionate. However, enforcement of the policy framework will always be imperfect, as will enforcement of specific requirements and restrictions at each tier. Some specific challenges remain:
Any technology permitting the personalisation of LLMs would need to comply with existing regulations and laws such as the GDPR , the Online Safety Bill and the various European AI standards . It is harder to assess whether an LLM is “fair” or “harmful” to its user when the space for possible personalisation actions is so complex. These issues also apply to LLMs in general, where adaptation of pre-trained models to downstream applications pose significant challenges to traditional auditing approaches .
It is challenging to regulate or evaluate a system which is both dynamic (adapting continually or frequently to user feedback) and distributed (personalised across many users). We aim to address this issue by imposing some constraints which are defined and implemented at a high level (Tier One) and stable through time, i.e., applying across all users and all training updates. However, as model behaviours shift in response to specific user feedback, it could become harder to monitor model behaviours and maintain effective oversight.
Our framework distributes responsibility among model providers, application providers and end-users, with the first two of these actors bearing the brunt of defining the bounds of personalisation. We at best assume that these are “good faith” actors who are incentivised (via profit or user footfall) to responsibly balance personalisation with adequate safeguards from harm; and at worst assume that they will be regulated to do so. In reality, with models like LLaMA being open-sourced, anyone with sufficient technical skill and compute resources could hypothetically create, launch and host a system capable of personalisation. Such fragmentation of the development landscape in LLMs poses a challenge for audit and policy enforcement .
While we define a clear functional framework, we have not fully specified principles that determine which restrictions or requirements fall under each tier. For example, why some risks come under Tier One (and are never allowed) while others are optionally defined in Tier Two. This is a problem also faced by online trust and safety regulation – where for example, early iterations of the Online Safety Bill included separate treatment of illegal content versus legal but harmful content, as well as tiered restrictions for children and minors versus adults. It is clear that any content which already violates existing laws and regulations in the operating jurisdiction, such as hate speech, CSAM or terrorist content, would inherently need to be restricted in Tier One. Beyond this, we leave the open question of what principles could be employed by model or application providers to define their own organisational bounds.
3 Next Steps
A large amount of our work in this paper, both in the taxonomy and in the policy framework, relies on informed speculation over the future development and governance of personalised LLMs. The affordances, constraints and harms from any technology depend critically on how it is designed, how its outputs are used in the real world and what safeguards or regulations are provisioned to guide its impact post-deployment. None of these conditions are presently clear; so, we very much consider this to be a first version of our work. In the future, we plan to iterate on the taxonomy and policy framework by conducting semi-structured interviews with (i) the end-users of LLMs to understand their priorities and concerns; (ii) model and application providers to scope their desires for enabling personalisation and for defining limits, as well as the technical apparatus available to them to both of these things; and finally, (iii) policymakers to assess how personalised LLMs fit into existing laws and regulations, and the new challenges they pose. By starting the conversation now, we hope to avoid long lags in understanding, documenting and governing the harms from personalised LLMs as a future technology which could widely impact individuals and society.
Acknowledgements
This paper, which began as an exploration of how to improve feedback between humans-and-model-in-the-loop, is part of a body of work funded by a MetaAI Dynabench grant. H.R.K’s PhD is supported by the Economic and Social Research Council grant ES/P000649/1. P.R’s PhD is supported by the German Academic Scholarship Foundation. We particularly want to thank Andrew Bean for his thoughtful input and assistance with the literature review. This paper is also the product of many interesting conversations at various conferences and working groups. We hope to survey the opinions and feedback of many other stakeholders in the future (including end-users, policy makers and technology providers) to further enrich our discussion.