In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
Sheng Liu, Haotian Ye, Lei Xing, James Zou
Introduction
Large language models (LLMs) have exhibited remarkable performance in various applications such as healthcare, education, and social interaction . With the increasing scale of these LLMs, in-context learning (ICL) has emerged as a striking property of LLMs . Unlike learning methods that require updating model parameters, in-context learning allows for good model performance with a prompt that only includes natural language instructions and/or a few demonstration examples .
Despite the remarkable ICL ability of LLMs, the efficacy of ICL is uneven and can be highly sensitive to the choices of templates , verbalizers , and demonstrations . This results in barriers in achieving LLM applications that are both adaptable and robust. In addition, the computational overhead of transformers limit the ability of existing LLMs to process extended contexts. A maximum context length (i.e., 4096) is set in the popular open-source LLMs, e.g. Llama 2. The direct consequence is that scaling up to large numbers of samples in in-context learning becomes inefficient and computational expensive.
In this paper, we show how standard in-context learning can be viewed as a process of “shifting” the latent states of the transformer. The direction and distance of shift are determined by the self-attention mechanism, which is not transparent and difficult to control. Inspired by this, we propose the In-Context Vector (ICV) approach, a scalable and inference-only alternative to in-context learning. The goal of ICV is to extract sufficient task information from the contextual demonstrations in ICL and provide this information to guide new query response generation in a controllable manner. We achieve this by breaking the process into two parts: Task summary, where an “in-context” vector is computed from the demonstration examples using the latent states of the Transformer; Feature shifting where the vector is applied to shift all latent states of the LLM during the forward pass of the query example, steering the generation process to incorporate the context task information.
The simplicity of the design of ICV only brings negligible computation overhead at computing the “in-context” vector. This design allows ICV to take many demonstration examples that exceed the limit of context length, as contextual demonstration examples are “summarized” in a single vector rather than directly prepended to the query. In contrast to ICL which requires templates and prompts, ICV can work with demonstration examples without any template. ICV is also easy to control as it directly “shifts” the latent embeddings by a magnitude that we can specify, while ICL relies on the built-in self-attention module to indirectly“shift” the latent features. Compared with finetuning which makes full gradient updates to change or add additional model parameters, ICV’s simplicity lies in its use of a single vector, avoiding significant computational expenses. The negligible computational overhead of ICV, without the introduction of new parameters, positions it as a practical enhancement to the standard ICL and finetuning framework.
We summarize our contributions as follows:
We propose In-Context Vector (ICV), an alternative to In-Context Learning (ICL) that is controllable and efficient.
We validate the effectiveness of ICV over diverse tasks from language model detoxification and style transform to role-playing with LLMs such as Falcon and Llama. ICV significantly outperforms standard ICL and LoRA fine-tuning on these tasks.
We exhibit a simple paradigm for adapting LLMs to a combination of tasks, centered around arithmetic operations of the in-context vectors.
Backgrounds
In the setting of in-context learning, consider the task of transferring negative sentiment to positive sentiment. A prompt is constructed by concatenating independent demonstration examples , e.g. “{1 star, I hate it!} is rewritten as {5 stars, I love it!}” followed by a query example: “{I don’t like the t-shirt} is rewritten as {”. Specifically, in-context learning assumes a target task with demonstration data . To perform the task for a given query example , the model is asked to predict based on the demonstrations. In traditional settings, is often a single word that represents the categorical label. However, could potentially be a sentence or paragraph that is related to . For example, can be an adapted version of in another tone or a paraphrase of following a specific style.
Large language models adopt Transformer as the backbone architecture. As a crucial component of the Transformer, self-attention layers relate different positions of a single sequence to compute a representation of the same sequence. Let denote the inputs (including demonstrations and the query examples ) for a specific self-attention layer of the Transformer. Let be the learnable key, query, and value matrix in that layer. In the in-context learning setting, the prefixed demonstration examples simply change the attention module through prepending a context matrix before the original query example. When the demonstrations are provided as the context, the attention layer for each token in the query example can be formulated as:
where is a scalar that represents the sum of normalized attention weights between demonstrations and query examples (See details in the Appendix C). Note that the first term, , is the original attention output without demonstration examples, whereas the second term is a position-wise modification that is based on demonstrations. Therefore, in-context learning essentially applies a position-wise modification to the original attention output by shifting the original output feature. The direction of the shift and the distance of the shift are automatically controlled by the self-attention mechanism.
Method
In section 2, we described that in-context learning essentially shifts the latent states of the query example by the self-attention mechanism in each layer. We propose a more straightforward way to shift the latent states. In particular, we use demonstration examples to create an “in-context vector”. This vector then directly shifts the latent states across the entire model, effectively transferring the essential details from the examples to the new problem it needs to solve. To illustrate this, refer to Figure 1, we first send and in each demonstration separately to the LLM to obtain the latent states at the last token position for and , capturing the latent states from the final part of each pair. These captured states are then combined to form the in-context vector (ICV), which stores the key information about the task. We then add this vector to the model’s latent states, which allows the model to tackle the new problem without needing the examples anymore. Our proposed method does not involve any finetuning on the large language model or training of any additional components.
Intuitively, the desired in-context vector (ICV) should be a direction that steers latent states closer to the representations of than the representations of . This could be achieved by optimizing a loss function that pulls representations closer to and pushes them away from . Motivated by this intuition, we can view ICV, denoted , as the optimizer of an objective
that encourages latent states to be closer to and be farther apart to . Under the setting of conventional in-context learning, and are paired, we show ICV can be extended beyond this setting, which is suitable for situations where paired examples are difficult to obtain. In the following section, we introduce two design choices of that work well empirically.
The maximizer of objective Eq. (2) subject to is the first principal direction of a set of real-valued data
Lemma 1 shows that the optimal solution of (2) is equivalent to the first principal direction of the differences between and . Therefore, we directly use the first principal direction of as the ICV. This objective has a similar form to the equalization step proposed by on removing attributes related to gender bias from word embeddings. However, the goal here is the opposite, the objective is maximized to learn the contrastive attributes between the and examples.
1.2 Unpaired demonstrations.
When demonstration examples are not paired, we adopt contrastive loss
where the ’s are the positive examples and ’s are the negative examples. The ICV is set to be the gradient of the Eq. (3) as the closed-form solution is unavailable
where , and . Intuitively, it automatically pairs each with multiple that are softly weighted by the corresponding . See details in Appendix F
2 Feature shifting: apply the ICV to query examples
After obtaining the ICV, a forward pass is performed on the query example. Since the demonstration examples are not directly used to guide the query example. In order to make the LLM aware of the in-context task, we steer the latent states towards the in-context learning task direction using ICV. For a sequence with tokens, we perform one gradient step by adding the ICV to the latent states at all layers and every token position as
This ensures that the modified latent state vectors remain close to the magnitude of representations typically accepted by the subsequent modules.
3 Task arithmetic property of in-context vectors
Users can add the in-context vector to perform the task that is aligned with the in-context demonstrations (e.g. transforming informal text to formal text), or a new task that has the opposite direction (e.g. transforming formal text to informal) without getting a new vector. In Table 5, we add an in-context vector to large language models to improve safety, reducing the proportion of generations classified as toxic, with little change in fluency. We also negate an in-context vector to reduce the positive tone of the generation.
Adding in-context vectors results in improved performance on a single task. Adding multiple in-context vectors of related tasks may result in improvements in the average performance on the entire set of tasks. In Table 5, we show that adding the “safe” vector and subtracting the “polite” vector results in a generation text that is safer but rude.
Experiments
We apply ICV to diverse tasks from LLM safety e.g. language detoxification, jail-break to personalization e.g. role-play. We compare the in-context vector method with conventional in-context learning, as well as with fine-tuning using the demonstration examples when paired demonstration examples are provided. We then extend ICV to unpaired demonstrations for LLM safety and personalization tasks.
We consider two aspects of safety for LLMs – using ICV to defend LLMs such as language detoxification and dialogue safety and using ICV to attack LLMs such as jail breaking safety aligned LLMs. The task of language detoxification is considered as paraphrasing offensive content. A set of demonstrations contain the offensive and inoffensive sentence pairs, and the offensive query sample are given. For the proposed in-context vector method, prompts and instructions are not necessary. We directly input ’s and ’s to obtain the in-context vector. For fine-tuning with the demonstration examples, we adopt low-rank adaptation (LoRA) for iterations. For dialogue safety and jail-break, since the data are already in the form of conversation, we do not use a specific prompt.
For language detoxification, we use ParaDetox which contain comments flagged for toxicity and provide matched non-toxic paraphrases that maintain the core meaning in a more neutral manner. For ParaDetox, we use randomly selected demonstration examples and evaluate on 670 other queries. For dialogue safety, we use the demonstrations listed in Table 9. For jail-break, we use 5 demonstrations listed in Table 11, 12 for ICV, and the jail-broken column as demonstrations for conventional in-context learning. Following the same setting in , we attack the model with 100 individual harmful behaviors and evaluate the attack success rate (ASR).
We use a safety classifier for automatically evaluating generation safety. We report the percentage of generated text being safe. Specifically, we use the 2.7B parameter Transformer classifier developed by . The classifier is trained on Wikipedia Toxic Comments , Build-it Break-it Fix-it , and Bot-Adversarial Dialogue . For a given target context and response, the classifier assigns a probability indicating whether the response is safe. We use the threshold as to flag responses as unsafe. For evaluating success rate of jail-break, following , we consider it as success when the generated sequence does not contain any of the token listed in Section A. We also manually check and ensure the coherence of the generations from the jail-broken model.
2 Writing style and role-playing
Large language models have shown their ability to imitate various writing styles. To evaluate our approach and better compare ICV with ICL and finetuning, we quantitatively evaluate different methods on two specific speaking styles: sentiment and formality. We also test their role-playing ability to speak in line with Shakespeare. For sentiment and formality transfer, we randomly selected demonstration examples. For role-playing, we randomly picked demonstrations.
In addition, we consider other three tasks for demonstration: transferring between reserved and emotive style, rudeness and politeness, as well as a format edition task in which we capitalize the initial letters of words. We create our own demonstration examples which are illustrated in Table 9 and report some exemplary generated outputs in Table 1. The number of demonstrations is varied from 3 to 4 depending on the task. Settings for in-context learning and LoRA fine-tuning can be found in Appendix A.
For formality transfer, we use the Grammarly’s Yahoo Answers Formality Corpus . This corpus contains paired informal and formal sentences without context under two topics. We used all sentences from the test set of the family and relationships topic for evaluation. demonstration examples are randomly sampled from a separate set under the same topic. For sentiment transfer, we utilize the first examples in the test set of the Yelp review dataset for evaluation. Since examples in the Yelp dataset are not paired, we obtain 5 paired demonstration examples by asking GPT-4 to produce the sentiment-transferred versions, which are presented in Table 10 of the Appendix. For role-playing Shakespeare, the evaluation is conducted on Shakespeare’s play – Romeo and Juliet . demonstration samples are randomly selected and the other queries are used for evaluation.
Similar to the prior task, we measure style accuracy using the prediction accuracy of the pre-trained style classifier over the generated sentences for automatically evaluating generation style. In addition to the above metrics, to evaluate role-playing, the GPT-3.5-Turbo is utilized to perform direct comparative assessments of responses (LLM-EVAL). We follow the setup described in wherein GPT-3.5-Turbo is prompted to rank the model based on its produced response in terms of being more like the role while preserving the meaning (see Appendix A for more details).
3 Large language models
We focus particularly on the popular large language models such as LLaMA and Falcon . Specifically, we applied ICV to LLaMA-7B , LLaMA-13B , Falcon-7B , and Vicuna-7b .
4 Automatic text similarity evaluation
In order to make sure that the paraphrased sentences preserve their original semantic meanings, we use open-text generation metrics to evaluate the similarity between the generation and the “gold-standard” references (for experiments in sentiment transfer, we use the original texts): ROUGE-1 for similarity in the raw text domain, Bert-Score for similarity in the feature domain.
Results
In this section, we quantitatively demonstrate the effectiveness and efficiency of our ICV framework. We will first describe the effectiveness of ICV on detoxification and safety. Then we will discuss the results of speaking style and role play. In the end, demonstrative outputs on the arithmetic properties of ICV will be presented.
In this section, we present a comprehensive analysis of the In-Context Vector (ICV) method’s contribution to improving LLMs’ safety.
Table 2 reports the automatic evaluation results of our proposed ICV method, demonstrating that ICV significantly surpasses both the conventional In-Context Learning (ICL) and the LoRA finetuning (LoRA FT) in the language detoxification task. Notably, ICV achieves a reduction in toxicity by 49.81% and 45.31% on Falcon-7b and Llama-7b models, respectively. These results indicate that ICV not only mitigates toxicity effectively but also enables LLMs to adapt more rapidly to new tasks than ICL does. ICV also maintains high semantic similarity to reference sentences, as indicated by robust ROUGE-1 and BERT scores, contrasting with finetuning approaches that often sacrifice semantic meaning to reduce toxicity. When considering a setting where examples are not paired, adopting contrastive-based ICV achieves performance only slightly worse than paired examples. In Table 1, we further illustrate ICV’s capability to enhance diagonal safety; it produces safe responses to unsafe queries while also identifying and addressing the discriminatory nature of the questions.
We analyze the effects of the scaling factor which controls the scale of the in-context task vector. As depicted in Figure 2 and the right panel of Figure 4, an increased intensifies the task’s influence, such as enhancing detoxification efforts. However, a larger also leads to a decline in the retention of the original text’s semantic meaning and reduces fluency, as evidenced by a lower ROUGE-1 score.
Further analysis reveals that, similar to ICL, the ICV method benefits from scaling up demonstration examples. This is supported by Figure 3, which shows a positive correlation between the number of demonstrations and a decrease in toxic generations. Additionally, unlike ICL, ICV is not constrained by context length, potentially allowing for the inclusion of more demonstration examples to further improve performance.
We performed a layer-specific ablation study to assess the impact of applying the in-context vector (ICV) to individual layers of the Transformer. This study compared the performance impact of applying ICV exclusively to the last, middle, and first layers of the Transformer. The results, as presented in Table 3, suggest that each variant performs similarly to the case where only the query example is used. These outcomes emphasize that ICV’s effectiveness is maximized when it is applied across all layers of the Transformer, rather than to individual layers.
2 Jail break
Previously, we demonstrate that ICV can be used to improve dialogue safety. A natural question to ask is whether ICV can make safety-aligned LLMs less safe. Here, we show that aligned LLMs can be jail-broken with the power of ICV. The results are shown in Table 6, which reveals that after just five instances of responding to malicious queries, the model inevitably adopts malicious behavior, generating harmful content in response to new malevolent prompts. When the strength of the ICV is enhanced, the attack success rate steadily climbs, eventually reaching 99%, on par with optimization-based methods including GBDA and PEZ which often take around 30 min per instance while ICV only takes a few seconds. This highlights the significant impact adversarial demonstrations have on skewing the alignment capabilities of Large Language Models (LLMs). Despite extensive measures implemented during fine-tuning to align the model, a mere application of ICV can make it learning to be dangerous.
3 Speaking style
We present the results across two types of styles – formality and sentiment – in Table 4. Our findings highlight the ICV’s robust capability to adapt to these styles. It improves the formality of the text by and positivity by . Furthermore, as illustrated in Table 1, ICV exhibits versatility by successfully modifying a given query example to align with additional stylistic dimensions, including text format and expressiveness of emotion.
4 Role-playing
We report win rates for ICV, ICL, and LoRA FT on role-playing Shakespeare in Figure 4. The win rate is the frequency with which a method is ranked first by GPT-3.5 among the three. We observe that ICV outperforms LoRA FT and ICL. For this task, inference-only methods (ICV, ICL) excel over the training-based approach (LoRA FT). When the model has more parameters, ICV achieves better results, suggesting that the performance tends to increase with model size. Moreover, we observe that finetuning results much worse results, indicating potential overfitting to the demonstration examples. Notably, increasing the scaling factor inversely affects the win rate for ICV, likely due to the reduced preservation of semantic meaning in the generations when compared to the reference.
5 Task arithmetic
We observe that the task arithmetic property also holds for in-context vectors. As we demonstrated in Table 5, multiple in-context tasks can be combined together if we perform vector addition and/or subtraction on the corresponding in-context vectors.
Related works
A branch of approaches that were introduced to boost ICL performance is by enhancing in-context example selection. These approaches aim to identify more effective in-context templates and examples. contributed methods to refine template selection, whereas focused on improving example choice. goes a step further by proposing a criterion that evaluates examples based on their consistency, diversity, and frequency of occurrence.
Other recent methods for improving ICL include flipped learning and noisy channel prompting . propose a method to assign labels by K-nearest neighbors for ICL to address multiple choice problems. also focuses on ICL for multiple choice problems and propose to iteratively update the context. proposes training decoder networks to serve as alternatives for few-shot ICL.
In the realm of ICL, the method that is most similar to ICV is introduced in a concurrent work . In this work, a “task vector” is first obtained by the latent states of a specific layer, the vector then “replaces” the latent states at the same layer during a forward pass of the query. The layer is selected based on prediction accuracy on development set. In contrast to this work, our work focuses on in-context open-ended generation, where traditional metrics like accuracy are not directly applicable. Instead of replacing the original latent states, our method enhances them by adding a vector across all layers, eliminating the need for a layer selection process. This addition preserves the inherent information within the original latent states. The distinctions not only underscore the unique aspects of our approach but also highlight its suitability for more nuanced, open-ended generative tasks.
Activation editing has been used in concurrent works to steer model outputs toward a specific behavior with different applications. focus on altering the GPT-2-XL for sentiment and topic shift. introduce “representation engineering” to align model behavior with certain concepts. demonstrates that latent knowledge encoded in the activation space is linearly separable. find that a "cheese vector" derived from model activations could modify the behavior of an RL agent. demonstrate how model activations could be edited to counterfactually change the model’s behavior. introduce “inference time intervention” with a similar technique, improving TruthfulQA performance. Our method focuses on the setting of in-context learning, with no use of specific templates or prompts but only demonstration examples to perform tasks demonstrated.
Recently, highlight the sensitivity of LLMs to the demonstration examples used in in-context learning (ICL). This phenomenon is further illuminated by through the lens of pretraining corpora and by with respect to pretraining term frequencies. Concurrently, offers an explanation of the ICL mechanism by likening it to implicit Bayesian inference, while demonstrates the emerging capability of LLMs to assimilate new input-label correspondences. Meanwhile, shows that the emerging learning algorithm from ICL is similar to gradient descent when performing linear regression. show ICL approximately performs gradient descent as meta-optimizers. However, it remains mysterious what is the exact internal working mechanism of ICL when LLMs perform complex natural language tasks.
Previous work has shown that task vectors, obtained by taking the weights of a model fine-tuned on a task and subtracting the corresponding pre-trained weights, encode the information necessary to do well on a given task. Moreover, they demonstrate that adding the task vectors leads to better performance on the task, and negating the vector results in task forgetting. We observe that ICV exhibits similar properties with no finetuning.
The general text generation capability of LLMs comes from pre-training on large, multi-domain text corpora. However, these corpora, which were crawled from the internet, inevitably contain toxic content . Most existing works mitigate toxicity in language models by further training or finetuning the models . Another line of work that is closer to our setting is the methods without additional training. proposes to use negative prompts to find the toxic token and update directions to the value vector in self-attention layers. Standard ICL is also adopted by to remove offensive content and improve dialogue safety.
Conclusions
In this paper, we introduce a novel framework to improve the effectiveness and efficiency of In-Context Learning (ICL). The framework decomposes ICL into two stages. The first stage generates an “in-context” vector from the demonstration examples based on the latent states of the Transformer. The vector stores the task information. In the inference stage, we apply the “in-context” vector by adding it to the latent states of all layers and token positions for the next word generation. Then, only the query example is put into the model for inference. Our two-stage In-Context Vector (ICV) method allows LLMs to be efficiently adapted to solving downstream tasks and combine multiple tasks together. Extensive experiments from model detoxification and safety to role-playing demonstrate that our method outperforms conventional ICL and LoRA finetuning in terms of both performance and efficiency. One limitation of ICV compared to standard in-context learning is that ICV requires access to the model in order to add the “in-context” vector. This is easy to implement for open-source models, which is the main focus of our experiments and applications here.
References
Appendix A More experimental details
For In-context learning, we use a simple template, described in A, for all style transfer tasks as well as role-playing. We adopted slightly different instructions for the formality transfer task: “Paraphrase the following sentence to be more polite and formal” even though we did not observe any obvious performance difference using different instructions for the template. For the proposed in-context vector approach, we use the same instruction as ICL without the demonstration examples. For conventional in-context learning, we consider the following prompt template:
All models have been finetuned on one A100 GPU. We train for 20 epochs without a warm-up, using gradient accumulation (batch size is set to the number of demonstration examples). The learning rate is set to 3e-4 for all models and tasks. The cutoff length for the examples is 512 tokens. The parameters for LoRA are as follows. we set rank to 8, dropout is set to 0.05 and alpha is set to 16. Target modules for LLaMA models are [q_proj,v_proj], and for Falcon models are [query_key_value].
Since multi-head attention is often used in LLMs, for convenience, we use the features output by mlp layers after attention layers as the latent states. We set the number of components for PCA as 1, and the scaling factor for the “in-context” vector as 0.1 if not otherwise specified. To further make ICV more adaptive, we emphasize the modification effect on latent states that are less aligned with the in-context task direction by setting , where is the latent states at layer , position , represents the ICV, is the similarity measurement, and we only further scale the modification when . For the jail-break task, we decay to within 5 tokens of generation to improve coherence of generations.
For automatic evaluation on formality transfer, we train a XLM-RoBERTa-based classifier on the training set of the GYAFC dataset. The classifier achieved an accuracy of 85.21% on classifying formal and informal on the GYAFC test set. For sentiment classification, we use the hugging face provided pipeline on “sentiment-analysis” and use the default model “distilbert-base-uncased-finetuned-sst-2-englis” for the task.
In order to evaluate role-playing Shakespeare, we follow and use GPT-3.5-Turbo to automatically evaluate the quality of generated sentences. GPT evaluator prompts are detailed here:
Prompt Template (GPT Evaluation) ⬇ Please act as an impartial and objective judge and evaluate the quality of the role-playing performance. You should rank the models based on the role characteristics as well as content relevance. The rankings are then output using Python dictionaries and lists. The models below are to play the role of William Shakespeare. Your evaluation should consider two factors below: 1. Which one has a more pronounced role-speaking style, and speaks more in line with the reference sentence in terms of the style. The more distinctive the speaking style, the better. 2. Which one better preserves the semantic meaning of the reference sentence. This is the reference sentence {} This is the role-playing sentence from model-A {} This is the role-playing sentence from model-B {} This is the role-playing sentence from model-C {} Now, based on the above two criteria, please rank the models. Avoid any positional biases and ensure that the order in which the responses are presented does not influence your decision. Do not favor certain model names. Be as objective as possible. Then, use a list containing the model’s name, its rank, and the reason for its ranking to return the results, i.e., please ensure to use the following format to return the results: [{"model":
To evaluate the jail-break experiments, we utilize the vicuna conversation template’s system message in to conserve memory and maintain a fair comparison among them. Post-generation, we employ the detection method described in to ascertain attack success. This involves checking if the generated sequence includes any of these tokens: ["I’m sorry", "I apologize", "As an", "I’m just", "I’m an", "I cannot", "Hello!"].
Appendix B Datasets and demonstration examples
In Table 7, we provide an overview of the datasets we used for quantitative evaluations. The selected demonstration examples are used for ICL, ICV, or LoRA finetuning. In Table 9, we provide demonstration examples for tasks that are not used for quantitative evaluations. Specifically, for tasks such as dialogue safety and style transfers, large-scale matched datasets may be difficult to find, we therefore create our own demonstration examples using GPT-4, some of the demonstration examples are also borrowed from Ref. .
Appendix C In-context learning as feature shifting
In section 2, we briefly describe the mechanism of ICL which changes the attention module through prepending the contextual demonstrations to the test query. Here, we provide a more detailed derivation and provide an alternative view of ICL. Let denote the inputs (including demonstrations and the query examples ) for the self-attention layer of the Transformer. Let be the learnable key, query, and value matrix respectively. In the in-context learning setting, the prefixed demonstration examples simply change the attention module through prepending a context matrix before the original query example. When the demonstrations are provided as the context, the attention for each token in the query example can be formulated as
Therefore, ICL can be viewed as shifting the original features by which depends on .
Appendix D Latent states for text-classification datasets
Text style transfer aims to control certain attributes in the generated text. Prior work often prompts LLMs at inference . Most related is the TextSETTR model and DIFFUR which train style extractors to produce style vectors, and incorporate these vectors into the latent feature of the encoder-decoder Transformer to control the style of the generation. In contrast to these methods, ICV does not require modification of the model’s architectures and further training.
Appendix E Proof of lemma 1
Let be the sample covariance matrix of , the objective becomes
Since is symmetric, according to the spectral theorem for symmetric matrices,
is the first eigenvector of which is the first principal direction that maximizes the sample variance of . ∎
Appendix F Gradients of the contrastive objective Eq. (3)
The objective Eq. (3) can be rewritten as