A Bayesian Framework for Information-Theoretic Probing
Tiago Pimentel, Ryan Cotterell
Introduction
Pimentel et al. (2020b) recently undertook an information-theoretic analysis of probing. They argue that probing may be viewed as approximating the mutual information between a linguistic property (e.g., part-of-speech tags) and a contextual representation (e.g., BERT). Counter-intuitively, however, due to the data-processing inequality, contextual representations contain exactly the same information about any task as the original sentence, under mild conditions. When viewed under this lens, the goal of probing is not inherently clear. One limitation of Pimentel et al.’s analysis is that it focuses on the mutual information (MI)—to be of practical application, their argument requires that a probe matches the true distribution according to which the data were generated in the limit of finite training data. In contrast, our paper formulates an information-theoretic framework that is compatible both with model misspecification and the finite data assumption.
In his seminal work, Shannon (1948) occupied himself with the limit of communication. Indeed, mutual information can be described as the theoretical limit (or upper-bound) of how much information can be extracted from one random variable about another. However, this limit is only achievable when one has full knowledge of these random variables, including the true probability distribution according to which they are distributed. In practice, we will not have access to such information and it may be difficult to approximate. It follows that any system with imperfect knowledge of the random variable’s true distribution will only be able to extract a subset of this information.
With this in mind, we propose and motivate an agent-based framework for measuring information. We term our quantity Bayesian mutual information and show that it generalises Shannon’s MI, holistically accounting for uncertainty within the Bayesian paradigm—measuring the amount of information a rational agent could extract from a random variable under partial knowledge of the true distribution. In addition to the definition, our paper provides many useful theoretical results. For instance, we prove that conditioning does not necessarily reduce Bayesian entropy and that Bayesian mutual information does not obey the data-processing inequality. We argue that these properties make our Bayesian framework ideal for an analysis of learned representations.
In the empirical portion of our paper, we investigate both part-of-speech tagging and dependency arc labelling. Moreover, our information-theoretic measure holistically captures a notion of ease of extraction, limiting the amount of data available to solve the task. Intuitively, Bayesian MI shows that high dimensional representations, such as BERT, actually hurt performance in the very low-resource scenario, making less information available to a Bayesian agent than a simple categorical distribution. This is because, when little data is available, these agents overfit to the evidence under weak priors. In the high-resource scenario of English, ALBERT dominates the curves, making more information available than other contextualised embedders. In short, Bayesian mutual information reconciles the probing literature with its frequently posed question: how much information can be extracted from these representations?
Background: Information Theory
Information theory (Shannon, 1948; Cover and Thomas, 2006) provides us with a number of tools to analyse data and their associated probability distributions—among which are the entropy and mutual information. These are traditionally defined according to a “true” probability distribution, i.e. or We use uppercase letters to denote random variables (, ), lowercase for their instances (, ), and calligraphic fonts for their domain space (, ). which may not be known, but dictates the behaviour of random variables and . The atomic unit of information theory is the surprisal, which is defined as follows:
Arguably, the most important information-theoretic definition is its expected value, termed entropy:
Finally, another important concept is the mutual information (MI) between two random variables
Unfortunately, information theory has a few properties which do not conform to our intuitions about the mechanics of information in machine learning:
A related question that arises is how to estimate information in scenarios where the true distribution is not known. For instance, what is the surprisal of a learning agent with a belief who encounters an instance ? The straightforward answer would be to use eq. 1—nonetheless, this agent does not know the true distribution . This agent’s surprisal is usually taken according to its belief:
Similarly, this agent’s entropy has been historically defined exclusively according to this belief distribution (Gallistel and King, 2011):
We term this the belief-entropy. We can further extend this to a belief mutual information:This definition is found in both the cognitive sciences (Gallistel and King, 2011; Fan, 2014; Sayood, 2018) as well as in active learning (Houlsby et al., 2011; Kirsch et al., 2019).
We note this definition is not grounded in the true distribution in any form. In fact, about the belief mutual information, Gallistel and King (2011) state: “the subjectivity that it implies is deeply unsettling […] the amount of information actually communicated is not an objective function of the signal from which the subject obtained it”.
A Bayesian Approach to Information
The primary motivation for this paper is developing a series of tools that help us overcome the limitations of traditional information theory as applied to machine learning. Specifically, probing representations requires a data-dependent information theory. We thus formulate analogues of surprisal, entropy and MI in terms of Bayesian agents—using a framework heavily inspired by Bayesian experimental design (Lindley, 1956). We then prove this framework does not suffer the same infelicities as standard information theory in this context.
Our discussions will focus on Bayesian agents, so we start by formally defining them.
A Bayesian agent is a parameterised probability distribution (or set of distributions) and a prior .In the case that the Bayesian agent has more than one distribution, we still only have a single prior without loss of generality. Indeed, separate priors for each distribution is a special case where the parameters are partitioned. In the case of , we could define . Given data , the Bayesian posterior over is
Analogously, the Bayesian belief is defined as the following posterior predictive distribution
Upon encountering an instance , and after seeing a collection of data , this agent’s posterior-predictive Bayesian surprisal will be:
where is a data-valued random variable; for notational succinctness, we omit this random variable for the rest of the paper. We further define the posterior-predictive Bayesian entropy:
As can be readily seen, the Bayesian entropy is the expected value of the Bayesian surprisal—with this expectation taken over the true distribution.We put this in contrast to eq. 5—which takes this expectation over the belief itself—since the instances are in practice encountered with this true frequency. This distinction has been explicitly noted before, by e.g. Bartlett (1953). In this sense, the Bayesian entropy is a cross-entropy rather than a standard entropy.
2 Bayesian Mutual Information
Defining Bayesian mutual information within our framework requires a bit more care. First, in contrast to surprisal and entropy, mutual information is a functional of two random variables. We will name the second random variable . To talk about mutual information, we will consider a Bayesian agent with a collection of at least two beliefs, e.g. . The second belief is conditional, but otherwise follows Definition 1.
Given a collection of data , and a Bayesian agent with a pair of beliefs and and a prior , the Bayesian mutual information (Bayesian MI) is defined as
There is an important distinction between the Bayesian and Shannon MI—Bayesian MI decomposes as the difference between two cross-entropies, as opposed to two entropies.Cross mutual information (XMI) has been used in several previous work such as (Pimentel et al., 2019, 2020b; Bugliarello et al., 2020; McAllester and Stratos, 2020; Torroba Hennigen et al., 2020; Fernandes et al., 2021; O’Connor and Andreas, 2021). In those works, though, it was usually interpreted as a computational approximation to the truth-MI (or to -information (Xu et al., 2020), which is discussed later in the paper). In this work, we highlight the Bayesian MI’s (and XMI’s) relevance as a generalisation of Shannon’s MI.
3 An Illustrative Example
For the sake of argument, we assume two independent categorical random variables and , both with classes and uniformly distributed.
This illustrates an important aspect of Bayesian MI: it is grounded on the true distribution.
4 Theoretical Properties
We now prove a few relevant theoretical properties about our framework. We show that Bayesian MI is symmetric if and only if the agent’s beliefs respect Bayes’ rule. Then, we discuss why it does not respect the data-processing inequality, and its connection to mutual information and to -information (Xu et al., 2020).
It is a well known result that Shannon’s MI is symmetric, i.e.
This means that the knowledge one can extract from random variable about is the same as the knowledge one can extract from about . This is not true in general for Bayesian MI; as we will show, information-theoretic symmetry and Bayes rule are tightly related. As such, we consider in this section a Bayesian agent with a set of beliefs . We call the agent consistent if it respects Bayes’ rule, i.e.
the following theorem characterises when we have symmetry.
An agent’s Bayesian mutual information is symmetric, i.e.
for all distributions if and only if the Bayesian agent is consistent.
4.2 No Data-Processing Inequality
Another classical result from information theory is the data processing inequality. This theorem states that processing a random variable can never add information, only reduce it
Although theoretically sound, this theorem is very unintuitive from a practical perspective—effectively, processing noisy data can make it more useful. In fact, representation learning is a subfield of machine learning devoted precisely to finding functions which can extract more “informative” representations from some input. One such example is BERT (Devlin et al., 2019), a large pre-trained language model which produces contextualised representations from sentential inputs. These representations provably contain the exact same information about any task as the original sentence (Pimentel et al., 2020b)—in practice, though, they are much more useful for downstream models.
The data processing inequality does not hold for Bayesian information, making it a more intuitive information-theoretic measure for probing; pre-trained representation extraction functions can increase MI from a Bayesian agent perspective.
The data processing inequality does not hold for Bayesian information, i.e.
4.3 Relation to Mutual Information
The relationship between Bayesian mutual information and Shannon MI is relevant for our discussion. As mentioned in the introduction, Shannon was concerned with the limits of communication when he defined his measure. We now put forward an intuitive theorem about Bayesian information; it is upper-bounded by the true MI under a weak assumption about the agent’s beliefs.
Assuming the agent’s belief has a smaller Kullback–Leibler (KL) divergence when compared to the true than the marginal of its beliefs over , i.e.
In other words, the information any agent can extract from a random variable about another variable is upper-bounded by the true MI. We now define a well-formed belief, which we will use to analyse the Bayesian MI’s convergence:
We say the belief of a Bayesian agent is well-formed if and only if the true distribution is a possible belief, i.e.
Given this definition, we prove the Bayesian mutual information converges to the true MI under well-defined conditions.
If we assume a Bayesian agent’s set of beliefs and prior are well-formed and meet the conditions of Bernstein–von Mises Theorem (pg. 339, Bickel and Doksum, 2001).In the case where is discrete and finite, the only requirement is , for all values of (Freedman, 1963). Then,
4.4 Relation to Variational Information
Variational (-) information (Xu et al., 2020) is a recent generalisation of mutual information. It extends MI to the case where a fixed family of distributions is considered; in which the true distribution may or not be.
Suppose random variable is distributed according to . Let be a variational family of distributions. Then, -entropy is defined
and -information is defined as
Assume a Bayesian agent’s beliefs and prior meet the conditions of Kleijn and van der Vaart (2012), who extend the Bernstein–von Mises Theorem to beliefs which are not well-formed. Further, let . Then,
A Framework for Incremental Probing
The proposed Bayesian framework for information allows us to take into account the amount of data we have for probing. Crucially, previous work Pimentel et al. (2020b) failed to adequately account for the observation of data. In doing so, they only analysed the limiting behaviour of information, under which the probing enterprise is not fully sensible—given unlimited data and computation, there is no point in using pre-trained functions. Indeed, the higher-level motivation of this work is to find an information-theoretic framework which serves machine learning, and under which the goal of probing is inherently clear. To that end, we propose a relatively simple experimental design. We compute Bayesian mutual information, which is a function of the amount of data, to create several learning curves.
1 Probes as Bayesian Agents
The overall trend in NLP is to train supervised probabilistic models on task-specific data. We believe probabilistic probes should analogously be modelled this way—leading to results compatible with our empirical intuitions. We thus define a probe agent as a Bayesian agent with the pair of beliefs and a prior . Any prior could be chosen for our probing agents. Nonetheless we have no a priori knowledge of how the representations should impact our prediction task. As such, our priors are such that the initial distributions and are identical. A logical conclusion, is that the prior Bayesian MI should be zero:
On the opposite extreme—i.e. given unlimited data—a well-formed belief will likely converge to the true distribution, yielding the same results as by Pimentel et al. (2020b). Complementarily, an ill-formed belief will converge to the -information:
The novelty of our framework lies in the explicit analysis of information under finite data. Bayesian agents are used here to measure a notion of information directly related to ease of extraction—i.e. how much information could be extracted from the representations by a naïve agent with no a priori knowledge about the task itself. In other words, we ask the question: given a specific dataset , how much information do the representations yield about this task? This value is only a subset of the true MI, being upper-bounded by it.
We focus our analysis on the amount of information a Bayesian agent can extract from the representations about the task. However, we could as easily analyse the Bayesian entropy instead. We believe, though, that the Bayesian MI is an inherently more intuitive value than the entropy. This is because mutual information puts the Bayesian entropy in perspective to a trivial baseline—how much uncertainty would there be without the representations. Furthermore, it has a much more interpretable value: with no data its value is zero, while at the limit it converges to the true mutual information. In this paper, we are concerned with its trajectory, i.e., how fast does the Bayesian MI go up?
2 Ease of Extraction and Previous Work
A few recent papers have tried to deal with probe complexity in a more nuanced way. Hewitt and Liang (2019) argue for the use of selectivity to control for probe complexity. Voita and Titov (2020) and Whitney et al. (2020) use, respectively, minimum description length (MDL) and surplus description length (SDL) to measure the size (in bits) of the probe model. Pimentel et al. (2020a) argues probe complexity and accuracy should be seen as a Pareto trade-off, and propose new metrics to measure probe complexity. All of these papers define ease of extraction in terms of properties of the probe, e.g., its complexity and size.
Experiments and ResultsOur code is available in https://www.github.com/rycolab/bayesian-mi.
We focus on part-of-speech (POS) tagging and dependency-arc labelling in our experiments. With this in mind, we make use of the universal dependencies (UD 2.6; Zeman et al., 2020); analysing the treebanks of four typologically diverse languages, namely: Basque, English, Marathi, and Turkish. As our object of analysis, we look at the contextual representations from ALBERT Lan et al. (2020), RoBERTa Liu et al. (2019) and BERT Devlin et al. (2019),We use the pre-trained models made available by the transformers library (Wolf et al., 2019). using as a baseline the non-contextual fastText Bojanowski et al. (2017) and random embeddings. Random embeddings are initialised at the type level and kept fixed during experiments.
2 Probe
Our experiments focus on Bayesian agents with multi-layer perceptron (MLP) beliefs:
As previously discussed, the Gaussian and Dirichlet priors on the parameters will cause these models to initially place a uniform distribution on the output classes—as such, they will have an initial Bayesian MI of zero. We then expose the probe agent to increasingly larger sets of data from the task. Unfortunately, the posterior of eq. 30 has no closed form solution, so we approximate it with the maximum-a-posteriori probability , where . We obtain this MAP estimate using the gradient descent method AdamW (Loshchilov and Hutter, 2019) with a cross-entropy loss and L2 norm regularisation.L2 weight decay regularisation is equivalent to a Gaussian prior on the parameter space (pg. 350, Bishop, 1995). The posterior predictive belief of eq. 31 has a closed-form solutionThis posterior predictive distribution is equivalent to Laplace smoothing (Jeffreys, 1939; Robert et al., 2009).
For both analysed tasks, we run 50 experiments with log-linearly increasing data sizes, from 1 instance to the whole language’s treebank. For each of these individual experiments, we sample an MLP probe configuration. This probe will have 0, 1, or 2 layers—where 0 layers means a linear probe—dropout between 0 and 0.5, and hidden size from 32 to 1024 (log distributed). We then use the same architecture to train a probe for each of our analysed representations, plotting their Pareto curves.
3 Discussion
Fig. 2 presents pareto curves for part-of-speech tagging. These curves convey a few interesting results. The first is the intuitive fact that information is much harder to extract with random embeddings, although with enough training data their results slowly converge to near the fastText ones—this can be seen most clearly in English. This matches our theoretical framework: the true mutual information between the target task and either fastText or random embeddings is the same, thus, if our beliefs are well-formed, the Bayesian MI should converge to this value, although with different speeds. The second result is that ALBERT makes information more easily extractable than either BERT or RoBERTa in English, and that multilingual BERT is roughly equally as informative as fastText under the finite data scenarios of the other analysed languages. Finally, the last result goes against one of the claims of Pimentel et al. (2020a), who in light of their flat Pareto curves for POS tagging claimed that we needed harder tasks for probing. One only needs harder tasks if their measure of complexity is not nuanced enough—as we see, even POS tagging is hard under the low-resource scenarios presented in our learning curves.
Fig. 3 presents results for dependency arc labelling. These learning curves also present interesting trends. While the POS tagging curves seem to be on the verge of convergence for English, Basque and Turkish, this is not the case for dependency arc labelling. This implies that, as expected, dependency arc labelling is either an inherently harder task, or that the representations encode the necessary information in a harder to extract manner. These results, also highlight the importance of an information-theoretic measure being able to capture negative information—as evinced in Fig. 4. For the low-data scenario, the BERToid models hurt performance, as opposed to helping. This is because high-dimensional representations, together with a weak prior, allow the agent to easily overfit to the little presented evidence. On the other hand, fastText does not present the same problem, having a positive Bayesian MI even in a low-data setting.
An Intuitive Decomposition
We now present some basic results about our framework which, although not strictly necessary for the present study, help motivate it. They also serve as a justification for our choice of cross-entropy when formalising Bayesian entropy. With this in mind, we analyse information from the perspective of a fully Bayesian agent with a well-formed belief.We make the same analysis from the perspective of an agent with an ill-formed belief in App. A A classic decomposition of the cross-entropy is the following:
We posit a new interpretation for this equality.
Let be a parameter-valued random variable. The entropy of a consistent Bayesian agent with well-formed beliefs decomposes as
In other words, the cross-entropy is composed of the sum between the entropy itself—i.e. the “true” information the data source provides, or its inherent uncertainty—and how much information the data provides about its distribution itself.
where we use prior predictive distributions, as opposed to posterior predictive ones. From this equation, we find that SDL is the information a dataset gives a Bayesian agent about its model parameters.
While closely related to one another, the Bayesian MI, MDL and SDL converge to different values in the limit of infinite dataset sizes:
Conclusion
In this paper we proposed an information-theoretic framework to analyse mutual information from the perspective of a Bayesian agent; we term this Bayesian mutual information. This framework has intuitive properties (at least from a machine learning perspective), which traditional information theory does not, for example: data can be informative, processing can help, and information can hurt. In the experimental portion of our paper, we use Bayesian mutual information to probe representations for both part-of-speech tagging and dependency arc labelling. We show that ALBERT is the most informative of the analysed representations in English; and high dimensional representations can provide negative information on low data scenarios.
Acknowledgements
We thank Adina Williams, Vincent Fortuin, Alex Immer, Lucas Torroba Hennigen and Mário Alvim for providing feedback in various stages of this paper. We also thank the anonymous reviewers for their valuable feedback in improving this paper.
Ethical Considerations
The authors foresee no ethical concerns with the research presented in this paper.
References
Appendix A Ill-formed Beliefs Loose Information
For the sake of argument, we now assume an agent with an ill-defined belief and a prior . We will show that such Bayesian agents loose information, meaning that they will not obtain as much information about their optimal parameters as if they had a well-formed belief.
Assume are the optimal parameters for a Bayesian agent with ill-formed, but consistent beliefs. The information this agent will receive about its optimal parameters is
This proof follows from the Bayesian MI definition, from this Bayesian agent having consistent beliefs, and from the fact that the cross-entropy is an upper-bound to the entropy, with equality only when both probability distributions are the same—which by definition is not possible
Appendix B Measures of Information
Several other measures of information have been proposed, among them are the entropy (DeGroot, 1962), the Rényi entropy (Rényi, 1961; Lenzi et al., 2000), Bayes vulnerability (Alvim et al., 2019), and the Determinantal Mutual Information (DMI; Kong, 2020). None of these take an agent’s belief into consideration, and so our analysis is orthogonal to them. The work most similar to ours, in this respect, is Clarkson et al.’s (2005) investigation of how belief impacts information leakage—and its extension, by Hamadou et al. (2010), to the Rényi min-entropy. Importantly, the results obtained by Clarkson et al. can be similarly derived using our framework.
Appendix C A Note on Empirical Limitations
Estimating the true MI between two random variables is known to be a hard problem for which several methods have been proposed (for a detailed review, see McAllester and Stratos, 2020)—estimating the Bayesian MI may be equally challenging. Given knowledge of and having access to samples from , the Bayesian MI can be trivially estimated using the Bayesian surprisal’s sample mean. On the other hand, in a setting such as active learning, where one (by definition) does not have access to the true distribution —only to the belief—the best approximation to the Bayesian MI may indeed be the belief-MI (used by Houlsby et al. 2011) or the Bayesian surprise (used by Storck et al. 1995 and Itti and Baldi 2006, 2009). Finally, approximating the Bayesian MI in the cognitive sciences may be an even harder problem than estimating the true MI, since it would require approximating both the belief of a specific agent and the true distribution of an event.
Appendix D Proof of Symmetric Bayesian Mutual Information, Theorem 1
An agent’s Bayesian mutual information is symmetric, i.e.
for all distributions if and only if the Bayesian agent is consistent.
We will first prove that if the Bayesian MI is symmetric for all true distributions , then the Bayesian agent is consistent (the if case). We then prove the inverse proposition (the only if case), completing this if and only if theorem’s proof.
We show this, by relying on specific distributions where is deterministic, putting all probability mass in a single point, i.e. .
As we can show this same result for any value of and , we conclude these agents must have consistent beliefs, i.e.
We now show that all consistent agents will have symmetric MI
Appendix E Proof of No Data Processing Inequality, Theorem 2
The data processing inequality does not hold for Bayesian information, i.e.
is, thus, a deterministic function of , where the function mean-centres random variable . Finally, we also define as a Bernoulli distributed random variable:
where is a sigmoid function. We can further define the distribution
We now define a Bayesian agent, which correctly knows the relationship between , and , i.e. with well-formed beliefs and , and with a prior —this agent does not know the true value of parameter though. For this Bayesian agent
When given , however, this agent does not need to know , since the data is already mean-centred (there are no unknown parameters in ). This Bayesian agent’s conditional entropy given is
This concludes the example that a deterministic (mean-centring) function can help this Bayesian agent.
where (1) becomes a strict inequality if the belief , i.e. if the prior does not place all probability mass in the true parameters . ∎
Appendix F Proof of Bayesian MI is Upper-bounded by the True MI, Theorem 3
Assuming the agent’s belief is tighter than the marginal of its beliefs over , i.e. than . We show
We start by noting that the difference between the true MI, and its Bayesian counterpart is equal to the difference between two KL-divergences
To prove this theorem, we need to show that one KL-divergence is smaller than the other, i.e.
We can show this with a bit of algebraic manipulation and an assumption about our Bayesian agent
In this equations, (1) relies on the log sum inequality, while (2) assumes the following inequality
This is equivalent to our assumption that this agent’s estimate of is tighter than if the agent marginalised its beliefs over . While not necessarily true, in practice, if is discrete and has a small cardinality , a simple Laplace smoothed estimate of is likely to result in this inequality. One could instead assume an agent which uses a Monte Carlo sampling approximation for estimating from . This would switch the inequality (2) for an approximation, and result in an expected lower bound instead
Appendix G Proof of the Convergence to Mutual Information, Theorem 4
If we assume a Bayesian agent’s set of beliefs and prior are well-formed and meet the conditions of Bernstein–von Mises Theorem (pg. 339, Bickel and Doksum, 2001). Then,
The Bernstein–von Mises Theorem only applies to well-formed beliefs, i.e. beliefs which can model the true probability distribution—a condition which is satisfied by our assumptions to this theorem. By this theorem—and under a number of other specified conditions, e.g. absolute continuity of the prior in a neighbourhood around and continuous positive density at (see pg. 141 in van der Vaart 2000 for the full set of conditions)—we have
Now, we apply the continuous mapping theorem to analyse the convergence of the Bayesian entropy
Appendix H Proof of the Convergence to 𝒱𝒱\mathcal{V}-information, Theorem 5
Assume a Bayesian agent’s beliefs and prior meet the conditions of Kleijn and van der Vaart (2012), who extend the Bernstein–von Mises Theorem to beliefs which are not well-formed. Further, let . Then,
Kleijn and van der Vaart (2012) extend the Bernstein–von Mises Theorem to ill-formed beliefs, showing that, under specific conditions for the Bayesian belief and priors, the predictive posterior distribution converges to
where is a unique set of parameters which minimises the KL-divergence between and the true distribution , i.e.
Given this convergence property, we can finish the proof similarly to the one for the well-formed belief:
where is defined as , and (1) relies on the continuous mapping theorem. We now conclude this proof:
Appendix I Proof of the Intuitive Decomposition, Theorem 6
Let be a parameter-valued random variable. The entropy of a consistent Bayesian agent with well-formed beliefs decomposes as
Note that under the Bayesian MI only the information about the true model parameters, i.e. , matters
where (1) relies on the fact that the true places all probability mass on the value . Using this result, we can show the Bayesian mutual information in eq. 73 is the same as the KL-divergence.