DCI-ES: An Extended Disentanglement Framework with Connections to Identifiability

Cian Eastwood, Andrei Liviu Nicolicioiu, Julius von Kügelgen, Armin Kekić, Frederik Träuble, Andrea Dittadi, Bernhard Schölkopf

Introduction

A primary goal of representation learning is to learn representations r(x)r(\bm{x}) of complex data x\bm{x} that “make it easier to extract useful information when building classifiers or other predictors” (Bengio et al. 2013). Disentangled representations, which aim to recover and separate (or, more formally, identify) the underlying factors of variation z\bm{z} that generate the data as x=g(z)\bm{x}=g(\bm{z}), are a promising step in this direction. In particular, it has been argued that such representations are not only interpretable (Kulkarni et al. 2015; Chen et al. 2016) but also make it easier to extract useful information for downstream tasks by recombining previously-learnt factors in novel ways (Lake et al. 2017).

While there is no single, widely-accepted definition, many evaluation protocols have been proposed to capture different notions of disentanglement based on the relationship between the learnt representation or code c=r(x)\bm{c}=r(\bm{x}) and the ground-truth data-generative factors z\bm{z} (Higgins et al. 2017; Eastwood & Williams 2018; Ridgeway & Mozer 2018; Kim & Mnih 2018; Chen et al. 2018; Suter et al. 2019; Shu et al. 2020). In particular, the metrics of Eastwood & Williams 2018—disentanglement (D), completeness (C) and informativeness (I)—estimate this relationship by learning a probe ff to predict z\bm{z} from c\bm{c} and can be used to relate many other notions of disentanglement (see Locatello et al. 2020).

In this work, we extend this DCI framework in several ways. Our main idea is that the functional capacity required to recover z\bm{z} from c\bm{c} is an important but thus-far neglected aspect of representation quality. For example, consider the case of recovering z\bm{z} from: (i) a noisy version thereof; (ii) raw, high-dimensional data (e.g. images); and (iii) a linearly-mixed version thereof, with each cic_{i} containing the same amount of information about each zjz_{j} (precise definition in section 6.1). The noisy version (i) will do quite well with just linear capacity, but is fundamentally limited by the noise corruption; the raw data (ii) will likely do quite poorly with linear capacity, but eventually outperform (i) given sufficient capacity; and the linearly-mixed version (iii) will perfectly recover z\bm{z} with just linear capacity, yet achieve the worst-possible disentanglement score of D ⁣=0D\!=0. Motivated by this observation, we introduce a measure of explicitness or ease-of-use based a representation’s loss-capacity curve (see fig. 1).

Structure and contributions. First, we connect the DCI metrics to two common notions of linear and nonlinear identifiability (section 3). Next, we propose an extended DCI-ES framework (section 4) in which we: (i) introduce two new complementary measures of representation quality—explicitness (E), derived from a representation’s loss-capacity curve, and size (S); and then (ii) elucidate a means to compute the D and C scores for arbitrary black-box probes (e.g., MLPs). Finally, in our experiments (section 6), we use our extended framework to compare different representations on the MPI3D-Real (Gondal et al. 2019) and Cars3D (Reed et al. 2015) datasets, illustrating the practical usefulness of our E score through its strong correlation with downstream performance.

Background

For step (ii), Eastwood & Williams 2018 use R\bm{R} and the prediction error to define and quantify three desiderata of disentangled representations: disentanglement (D), completeness (C), and informativeness (I).

Disentanglement. Disentanglement (D) measures the average number of data-generating factors zjz_{j} that are captured by any single code cic_{i}. The score DiD_{i} is given by Di=1−HK(Pi.){D_{i}=1-H_{K}(P_{i.})}, where HK(Pi.)=−∑k=1KPiklog⁡KPikH_{K}(P_{i.})=-\sum_{k=1}^{K}P_{ik}\log_{K}P_{ik} denotes the entropy of the distribution Pi.P_{i.} over row ii of RR, with Pij=Rij/∑k=1KRikP_{ij}=R_{ij}/\sum_{k=1}^{K}R_{ik}. If cic_{i} is only important for predicting a single zjz_{j}, we get a perfect score of Di=1D_{i}=1. If cic_{i} is equally important for predicting all zjz_{j} (for j ⁣= ⁣1,…,Kj\!=\!1,\dots,K), we get the worst score of Di=0D_{i}=0. The overall score DD is then given by the weighted average D=∑i=1LρiDiD=\sum_{i=1}^{L}\rho_{i}D_{i}, with ρi=1K∑k=1KRik\rho_{i}=\frac{1}{K}\sum_{k=1}^{K}R_{ik}.

Together, D and C quantify the degree of “mixing” between c\bm{c} and z\bm{z}, i.e., the deviation from a one-to-one mapping. They are reported separately as they capture distinct criteria.

Connection to identifiability

The goal of learning a data representation which recovers the underlying data-generating factors is closely related to blind source separation and independent component analysis (ICA, Comon 1994; Hyvärinen & Pajunen 1999; Hyvarinen et al. 2019). Whether a given learning algorithm provably achieves this goal up to acceptable ambiguities, subject to certain assumptions on the data-generating process, is typically formalised using the notion of identifiability. Two common types of identifiability for linear and nonlinear settings, respectively, are the following.

We say that c=r(x)=r(g(z))\bm{c}=r(\bm{x})=r(g(\bm{z})) identifies z\bm{z} up to sign and permutation if c=Pz\bm{c}=\bm{P}\bm{z} for some signed permutation matrix P\bm{P} (i.e., ∣P∣|\bm{P}| is a permutation).

We say c\bm{c} identifies z\bm{z} up to permutation and element-wise reparametrisation if there exists a permutation π\pi of {1,...,K}\{1,...,K\} and invertible scalar-functions {hk}k=1K\{h_{k}\}_{k=1}^{K} s.t. ∀j\forall j: cj=hj(zπ(j))c_{j}=h_{j}(z_{\pi(j)}).

We now establish theoretical connections between the DCI framework and these identifiability types.

If D=C=1D=C=1 and K=LK=L (i.e., dim(c)=dim(z)\text{dim}(\mathbf{c})=\text{dim}(\mathbf{z})), then R\bm{R} is a permutation matrix.

All proofs are provided in appendix A. Using proposition 3.3, we can establish links to identifiability, provided the inferred representation c\bm{c} perfectly predicts the true data-generating factors z\bm{z}, i.e., I=1I=1.

Under the same conditions as proposition 3.3, if z=W⊤c\bm{z}=\bm{W}^{\top}\bm{c} (so that I=1I=1) for some W\bm{W} with Rij=∣wij∣∑i=1L∣wij∣R_{ij}=\frac{|w_{ij}|}{\sum_{i=1}^{L}|w_{ij}|}, then c\bm{c} identifies z\bm{z} up to permutation and sign (definition 3.1).

For nonlinear ff, we give a more general statement for suitably-chosen feature-importance matrices R\bm{R}.

Under the same conditions as proposition 3.3, let z=f(c)\bm{z}=f(\bm{c}) (so that I=1I=1) with ff an invertible and differentiable nonlinear function, and let R\bm{R} be a matrix of relative feature importances for ff (definition 2.1) with the property that Rij=0R_{ij}=0 if and only if fjf_{j} does not depend on cic_{i}, i.e., ∣∣∂ifj∣∣2=0{\left|\left|\partial_{i}f_{j}\right|\right|}_{2}=0. Then c\bm{c} identifies z\bm{z} up to permutation and element-wise reparametrisation (definition 3.2).

While the if part of corollary 3.5 holds for most feature importance measures, the only if part, in general, does not: not using a feature cic_{i} is typically a sufficient condition for Rij=0R_{ij}=0, but it need not be a necessary condition (as required for corollary 3.5). E.g., measures based on average performance may not satisfy this since a feature may not contribute on average, but still be used—sometimes helping and sometimes hurting performance (see section 7 for further discussion). In contrast, Gini importances, as used in random forests, do satisfy the necessary condition. While the non-invertibility of random forests prevents an explicit link to identifiability (typically studied for continuous features), they can still be a principled choice in practice (where features are often categorical).

Summary. We have established that the learnt representation c\bm{c} identifies the ground-truth z\bm{z} up to:

sign and permutation if D ⁣= ⁣C ⁣= ⁣I ⁣= ⁣1D\!=\!C\!=\!I\!=\!1 and ff is linear;

permutation and element-wise reparametrisation if D ⁣= ⁣C ⁣= ⁣I ⁣= ⁣1 and Rij=0⇔∣∣∂ifj∣∣2=0D\!=\!C\!=\!I\!=\!1\text{ and }R_{ij}=0\Leftrightarrow{\left|\left|\partial_{i}f_{j}\right|\right|}_{2}=0.

Extended DCI-ES Framework

Motivated by our theoretical insights from section 3—considering different probe function classes provides links to different types of identifiability—and the empirically-observed performance differences between representations trained with different-capacity probes shown in fig. 1, we now propose several extensions of the DCI framework.

We first introduce a new complementary notion of disentanglement based on the functional capacity required to recover or predict z\bm{z} from c\bm{c}. The key idea is to measure the explicitness or ease-of-use (E) of a representation using its loss-capacity curve.

Loss-capacity curves. A loss-capacity curve for representation c\bm{c}, factor zjz_{j}, and probe class F\mathcal{F} displays test-set loss against probe capacity for increasing-capacity probes f∈Ff\in\mathcal{F} (see fig. 1). To plot such a curve, we must train TT predictors with capacities κ1,…,κT\kappa_{1},\dots,\kappa_{T} to predict zjz_{j}, with

Here κ1,…,κT\kappa_{1},\dots,\kappa_{T} is a list of TT increasing probe capacities, ideally True for RFs but not input-size dependent MLPs (see section 6). shared by all representations, with suitable choices for κ1\kappa_{1} and κT\kappa_{T} depending on both F\mathcal{F} and the dataset. For example, we may choose κT\kappa_{T} to be large enough for all representations to achieve their lowest loss and, for random forest ffs, we may choose an initial tree depth of κ1=1\kappa_{1}=1 and then T−2T-2 tree depths between 11 and κT\kappa_{T}.

Explicitness. We define the explicitness (E) of representation c\bm{c} for predicting factor zjz_{j} with predictor class F\mathcal{F} as

A fine-grained picture of identifiability. Compared to the commonly-used mean correlation coefficient (MCC) or Amari distance (Amari et al. 1996; Yang & Amari 1997), the D,C,I,ED,C,I,E scores represent empirical measures which: (i) easily extend to mismatches in dimensionalities, i.e., L>KL>K; and (ii) provide a more fine-grained picture of identifiability (violations), for if the initial probe capacity κ1\kappa_{1} is linear and R\bm{R} satisfies corollary 3.5, we have that:

D ⁣= ⁣C ⁣= ⁣I ⁣= ⁣E ⁣= ⁣1  ⟹  D\!=\!C\!=\!I\!=\!E\!=\!1\implies identified up to sign and permutation (definition 3.1);

D ⁣= ⁣C ⁣= ⁣I ⁣= ⁣1  ⟹  D\!=\!C\!=\!I\!=\!1\implies identified up to permutation and element-wise reparametrisation (definition 3.2);

I ⁣= ⁣E ⁣= ⁣1 ⁣  ⟹   ⁣I\!=\!E\!=\!1\!\implies\! identified up to invertible linear transformation (Khemakhem et al. 2020, cf.).

Thus, if D ⁣= ⁣C ⁣= ⁣I ⁣= ⁣E ⁣= ⁣1D\!=\!C\!=\!I\!=\!E\!=\!1 does not hold exactly, which score deviates the most from 11 may provide valuable insight into the type of identifiability violation.

Probe classes. As emphasized above, whether or not a representation c\bm{c} is explicit or easy-to-use for predicting factor zjz_{j} depends on the class of probe F\mathcal{F} used, e.g., MLPs or RFs. More generally, the explicitness of a representation depends on the way in which it is used in downstream applications, with different downstream uses or probe classes resulting in different definitions of explicit or easy-to-use information. We thus conduct experiments with different probe classes in section 6.

2 Size (S)

We next introduce a measure of representation size (S), motivated by the observation that larger representations tend to be both more informative and more explicit (see fig. 4, more details below). Reporting S thus allows size-informativeness and size-explicitness trade-offs to be analysed.

A measure of size. We measure representation size (S) relative to the ground-truth as:

When L ⁣≥ ⁣KL\!\geq\!K, as often the case, we have S ⁣∈ ⁣(0,1]S\!\in\!(0,1] with the perfect score being S ⁣= ⁣1S\!=\!1. However, if we also consider the L<KL<K case, which would likely sacrifice some informativeness, we have S ⁣∈ ⁣(1,K]S\!\in\!(1,K].

Larger representations are often more informative. When L ⁣< ⁣KL\!<\!K, it is intuitive that larger representations are more informative—they can simply preserve more information about z\bm{z}. When L ⁣> ⁣KL\!>\!K, however, it is also common for larger representations to be more informative, perhaps due to an easier optimization landscape (Frankle & Carbin 2019; Golubeva et al. 2021). fig. 4 illustrates this point, where AE-5 denotes an autoencoder with L ⁣= ⁣5L\!=\!5. Note that K ⁣= ⁣7K\!=\!7 for MPI3D-Real (see section 6).

Larger representations are often more explicit. The explicitness of a representation also depends on its size: larger representations tend to be more explicit, as is apparent from the second column of fig. 4. To explain this, we plot the corresponding loss-capacity curves in fig. 4. Here we see that the increased explicitness (i.e., smaller AULLC) of larger representations stems from a substantially lower initial loss when using a linear-capacity MLP probe. The fact that larger representations perform better with linear-capacity MLPs is unsurprising since they have more parameters.

3 Probe-agnostic feature importances

Finally, to meaningfully discuss more flexible probe-function choices within the DCI-ES framework, we point out that the D and C scores can be computed for arbitrary black-box probes ff by using probe-agnostic feature-importance measures. In particular, in our experiments (section 6), we use SAGE (Covert et al. 2020) which summarises each feature’s importance based on its contribution to predictive performance, making use of Shapley values (Shapley 1953) to account for complex feature interactions. Such probe-agnostic measures allow the DD and CC scores to be computed for probes with no inherent or built-in notion of feature importance (e.g., MLPs), thereby generalising the Lasso and RF examples of Eastwood & Williams 2018. While SAGE has several practical advantages over other probe-agnostic methods (see, e.g., Covert et al. 2020, Table 1), it may not satisfy the conditions required to link the DD and CC scores to different identifiability equivalence classes (see remark 3.6). Future work may explore alternative methods which do, e.g., by looking at a feature’s mean absolute attribution value (Lundberg & Lee 2017) since, intuitively, absolute contributions do not allow for a cancellation of positive and negative attribution on average (cf. remark 3.6).

Related work

Explicit representations. Eastwood & Williams 2018 noted that the informativeness score with a linear probe quantifies the amount of information in c\bm{c} about z\bm{z} that is “explicitly represented”, while Ridgeway & Mozer 2018 proposed a measure of “explicitness” which simply reports the informativeness score with a linear probe. In contrast, our DCI-ES framework differentiates between the amount of information in c\bm{c} about z\bm{z} (informativeness) and the ease-of-use of this information (explicitness). This allows a more fine-grained analysis of the relationship between c\bm{c} and z\bm{z}, both theoretically (distinguishing between more identifiability equivalence classes; section 3) and empirically (section 6).

Loss-capacity curves. Plotting loss against model complexity or capacity has long been used in statistical learning theory, e.g., for studying the bias-variance trade-off (Hastie et al. 2009, Fig. 7.1). More recently, such loss-capacity curves have been used to study the double-descent phenomenon of neural networks (Belkin et al. 2019; Nakkiran et al. 2021) as well as the scaling laws of large language models (Kaplan et al. 2020). However, they have yet to be used for assessing the quality or explicitness of representations.

Loss-data curves. Whitney et al. 2020 use loss-data curves, which plot loss against dataset size, to assess representations. They measure the quality of a representation by the sample complexity of learning probes that achieve low loss on a task of interest. Loss-data curves are also studied under the term learning curves in standard/purely supervised-learning settings (see, e.g., Viering & Loog 2021, for a recent review). In contrast, we focus on functional complexity and the task of predicting the data-generative factors z\bm{z}, and then discuss the functional complexity for other tasks y\bm{y} in section 7.

Experiments

Data. We perform our analysis of loss-capacity curves on the MPI3D-Real (Gondal et al. 2019) and Cars3D (Reed et al. 2015) datasets. MPI3D-Real contains ≈1\approx 1M real-world images of a robotic arm holding different objects with seven annotated ground-truth factors: object colour (66), object shape (66), object size (22), camera height (33), background colour (33) and two degrees of rotations of the arm (40×4040\times 40); numbers in brackets indicate the number of possible values for each factor. Cars3D contains ≈17.5k\approx 17.5k rendered images of cars with three annotated ground-truth factors: camera elevation (4), azimuth (24) and car type (183).

Representations. We use the following synthetic baselines and standard models as representations:

Noisy labels: c=z+ϵ\bm{c}=\bm{z}+\bm{\epsilon}, with ϵ∼N(0,0.01⋅IK)\bm{\epsilon}\sim\mathcal{N}(\bm{0},0.01\cdot\bm{I}_{K}).

Linearly-mixed labels: c=Wz\bm{c}=W\bm{z}, with Wij=1LK+ϵijW_{ij}=\frac{1}{LK}+\epsilon_{ij} and ϵij∼N(0,0.001)\epsilon_{ij}{\sim}\mathcal{N}(0,0.001) to achieve “uniform mixing” (each zjz_{j} evenly-distributed across the cic_{i}s) while also ensuring the invertibility of WW a.s.

Raw data (pixels): c=x=g(z)\bm{c}=\bm{x}=g(\bm{z}).

Others: We also use VAEs (Kingma & Welling 2014) with 1010 latents (L=10L{=}10), β\beta-VAEs (Higgins et al. 2017, L=10L{=}10); and an ImageNet-pretrained ResNet18 (He et al. 2016, L=512L{=}512).

Probes. We use MLPs, RFs and Random Fourier Features (RFFs, Rahimi & Recht 2007) to predict z\bm{z} from c\bm{c}, with RFFs having a linear classifier on top. For MLPs, we start with linear probes (no hidden layers) then increase capacity by adding two hidden layers and varying their widths from 2×K2\times K to 512×K512\times K. We then measure capacity based on the number of “extra” parameters beyond that of the linear probe, and compute feature importances using SAGE with permutation-sampling estimators and marginal sampling of masked values (see https://github.com/iancovert/sage). For RFs, we use ensembles of 100100 trees, control capacity by varying the maximum depth between 11 and 3232, and compute feature importances using Gini importance. For RFFs, we control capacity by exponentially increasing the number of random features from 242^{4} to 2172^{17} , and compute feature importances using SAGE.

2 Evaluation results: curves and scores

Loss-capacity curves. Figure 1 depicts loss-capacity curves for the three probes and two datasets, averaged over factors zjz_{j}. In all six plots, the noisy-labels baseline performs well with low-capacity and then is surpassed by other representations given sufficient capacity, as expected. Note that the linearly-mixed-labels baseline immediately achieves ≈0\approx 0 loss with MLP probes but not with RFF or RF probes, supporting the idea that the explicitness or ease-of-use of a representation depends on the way in which it is used. Also note that, with MLP probes and log⁡(excess #params)\log(\text{excess}\ \#\text{params}) as the capacity measure, larger input representations are afforded more parameters with a linear probe and thus are more expressive. This further explains why larger representations are often more explicit, and highlights the difficulty of measuring the capacity of MLPs—an active area of research in its own right, which we discuss in section 7. Finally, in section B.2, we investigate the effect of dataset size by plotting loss-capacity curves for different dataset sizes, observing that larger datasets have smaller performance gaps between: (i) synthetic and learned representations; and (ii) small and large representations (see fig. 11).

DCI-ES scores. Table 1 reports the corresponding DCI-ES scores, along with some oracle scores for MLPs. Note that: (i) the GT labels z\bm{z} get perfect scores of 11 for all metrics; (ii) by attaining very low D and C scores but near-perfect E scores, the linearly-mixed labels expose the key difference between mixing-based (D,C) and functional-capacity-based (E) measures of the simplicity of the c\bm{c}-z\bm{z} relationship; (iii) larger representations (ImgNet-pretr, raw data) tend to be more explicit than smaller ones (VAE, β\beta-VAE), with SS and EE together capturing this size-explicitness trade-off; and (iv) β\beta-VAE achieves better mixing-based scores (D,C) but similar E scores compared to the VAE, illustrating that these two “disentanglement” notions are indeed orthogonal and complementary.

3 Downstream results: score correlations

Analysis. Figures 5(a) and 5(b) show that E is strongly correlated with downstream performance when using both MLP (ρ ⁣= ⁣0.96,p ⁣= ⁣8e-18\rho\!=\!0.96,p\!=\!\text{8e-18}) and RF probes (ρ ⁣= ⁣0.88,p ⁣= ⁣2e-10\rho\!=\!0.88,p\!=\!\text{2e-10}). In contrast, mixing-based disentanglement scores (D, C) exhibit much weaker correlations with MLP probes, corroborating the results of Träuble et al. 2022 who also found a weak correlation between D and downstream performance on reinforcement learning tasks with MLPs. See App. B.1 for further details and results.

Discussion

Why connect disentanglement and identifiability? Connecting prediction-based evaluation in the disentanglement literature to the more theoretical notion of identifiability has several benefits. Firstly, it provides a concrete link between two often-separate communities. Secondly, it endows the often empirically-driven or practice-focused disentanglement metrics with a solid and well-studied theoretical foundation. Thirdly, compared to the commonly-used MCC or Amari distance, it provides the ICA or identifiability community with more fine-grained empirical measures, as discussed in section 4.1.

Measuring probe capacity. Our measure of explicitness EE depends strongly on the choice of capacity measure for a probe or function class. For some probes like RFs or RFFs, there exist natural measures of capacity. However, for other probes like MLPs, coming up with a good capacity measure is itself an important and active area of research (Jiang et al. 2020; Dziugaite et al. 2020). Another difficulty arises from choosing a capacity scale, with different scales (e.g., log, linear, etc.) leading to loss-capacity curves with different shapes, areas and thus explicitness scores. To investigate the extent of this issue, i.e., the sensitivity of our explicitness measure to the choice of capacity scale, fig. 6 compares the explicitness scores when using logarithmic and linear scaling. Here we see that the ranking essentially remains the same except for the raw-data representation with MLP probes.

Measuring feature importance. Similarly, the choice of feature-importance measure has a strong influence on the DD and CC scores, with some probes having natural or in-built measures (e.g., random forests) and others not (e.g., MLPs). For the latter, we proposed the use of probe-agnostic feature-importance measures like SAGE, and specified the conditions (corollary 3.5) that importance measures must satisfy if the resulting DD and CC scores are to be connected to identifiability. As with probe capacity, coming up with good measures of feature importance is its own orthogonal field of study (e.g., model explainability), with future advances likely to improve the DCI-ES framework.

What about explicitness for other tasks y\bm{y}? While we focused on the explicitness or ease-of-use of a representation for predicting the data-generative factors z\bm{z}, one may also be interested in its ease-of-use for other tasks/labels y\bm{y}. While it is often implicitly assumed that the ease-of-use for predicting z\bm{z} correlates with the ease-of-use for common tasks of interest (e.g., object classification, segmentation, etc.), future work could directly evaluate the explicitness of a representation for particular tasks y\bm{y}. For example, one could consider the entire loss-capacity curve when benchmarking self-supervised representations on ImageNet, rather than just linear-probe performance (a single slice). Future work could also explore the trade-off between explicit but task-specific and implicit but task-agnostic representations.

Conclusion

We have presented DCI-ES—an extended disentanglement framework with two new complementary measures of representation quality—and proven its connections to identifiability. In particular, we have advocated for additionally measuring the explicitness (E) of a representation by the functional capacity required to use it, and proposed to quantify this explicitness using a representation’s loss-capacity curve. Together with the size (S) of a representation, we believe that our extended DCI-ES framework allows for a more fine-grained and nuanced benchmarking of representation quality.

The authors would like to thank Chris Williams, Francesco Locatello, Nasim Rahaman, Sidak Singh and Yash Sharma for helpful discussions and comments. This work was supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A, 01IS18039B; and by the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645.

References

Appendix A Proofs

Now for any p=(p1,...,pK)∈ΔK−1\bm{p}=(p_{1},...,p_{K})\in\Delta_{K-1}, we have that

A.2 Proof of corollary 3.4

First, we show that Rij=∣wij∣∑i=1L∣wij∣R_{ij}=\frac{|w_{ij}|}{\sum_{i=1}^{L}|w_{ij}|} is a well-defined feature importance matrix. Suppose for a contradiction, that ∑i=1L∣wil∣=0\sum_{i=1}^{L}|w_{il}|=0 for some ll. Since ∣wil∣≥0|w_{il}|\geq 0, this implies wil=0w_{il}=0 for all ii. Consider zl=∑i=1Lwilciz_{l}=\sum_{i=1}^{L}w_{il}c_{i}. Taking the covariance, we obtain Var⁡[zl]=∑i,j=1LwilwjlCov⁡(ci,cj)=0\operatorname{Var}[z_{l}]=\sum_{i,j=1}^{L}w_{il}w_{jl}\operatorname{Cov}(c_{i},c_{j})=0, which is a contradiction since zlz_{l} has positive (unit) variance by the normalisation assumption (see footnote 1). Hence, ∑i=1L∣wil∣>0\sum_{i=1}^{L}|w_{il}|>0 for all ll. Thus R\bm{R} is well-defined, with its elements being non-negative and its columns summing to one by construction, so it is a valid feature importance matrix.

Next, note that we can write R=∣W∣D\bm{R}=|\bm{W}|\bm{D} where D\bm{D} is the invertible diagonal matrix with positive diagonal entries Djj=1∑i=1L∣wij∣>0D_{jj}=\frac{1}{\sum_{i=1}^{L}|w_{ij}|}>0.

By proposition 3.3, R\bm{R} is a permutation matrix, so R=P=∣W∣D\bm{R}=\bm{P}=|\bm{W}|\bm{D} for some permutation matrix P\bm{P}. Right multiplication by D−1\bm{D}^{-1} yields PD−1=∣W∣\bm{P}\bm{D}^{-1}=|\bm{W}|, that is ∣W∣|\bm{W}| has exactly one non-zero, positive element in each row and each column (and zeros elsewhere). Thus W\bm{W} and therefore also W⊤\bm{W}^{\top} are generalised permutation matrices. Hence (W⊤)−1(\bm{W}^{\top})^{-1} exists and is also a generalised permutation matrix.

A.3 Proof of corollary 3.5

For any jj consider zj=fj(c)z_{j}=f_{j}(\bm{c}). By proposition 3.3, RR is a permutation matrix, so column jj of R\bm{R} contains exactly one non-zero entry in row π(j)\pi(j) for some permutation π\pi of {1,...,K}\{1,...,K\}. Hence, by the assumed property of RR, fj(c)f_{j}(\bm{c}) does not depend on cic_{i} for all i≠π(j)i\neq\pi(j), and thus zj=fj(cπ(j))z_{j}=f_{j}(c_{\pi(j)}) ∀j\forall j. By invertibility of ff, we obtain cj=hj(zj′)c_{j}=h_{j}(z_{j^{\prime}}) with hj=fj′−1h_{j}=f_{j^{\prime}}^{-1} and j′=π−1(j)j^{\prime}=\pi^{-1}(j). ∎

Appendix B Additional Experimental Results

Here we present the full results of the correlations between the DCIE scores and downstream performance, the latter with low-capacity probes (as discussed in section 6.3).

In table 2 and table 3 we show the values of the Pearson and Spearman correlations alongside the corresponding pp-values Computed using https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.pearsonr.html. Note that some of the assumptions behind these pp-values, e.g. that the DCIE scores and downstream performances are normally distributed, likely do not hold. Thus, these pp-values should not be interpreted as precise probabilities but rather as rough indications of statistical significance. In table 4 we show the correlations for regression and classification tasks separately, with both task types exhibiting similar correlations. We note that E has the strongest correlation with the downstream performance (when using low-capacity probes for the downstream task).

To get a deeper insight into the correlations reported in tables 2 and 3, we plot each of the D, C, I and E scores against downstream performance for each of the 30 models considered in section 6.3. As shown in Figures 8, 9, 10 and 7, only E correlates strongly with downstream performance for both probe types, again highlighting: (i) the value that E adds to the existing DCI framework; and (ii) the practical usefulness of reporting E when comparing/evaluating learned representations.

B.2 Different amounts of data

In fig. 11 we present loss-capacity curves obtained when using different amounts of data to train the MLP probes. As shown, larger datasets have smaller performance gaps between (i) synthetic and learned representations; and (ii) small and large representations.