Independent mechanism analysis, a new concept?
Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf, Michel Besserve
Introduction
One of the goals of unsupervised learning is to uncover properties of the data generating process, such as latent structures giving rise to the observed data. Identifiability formalises this desideratum: under suitable assumptions, a model learnt from observations should match the ground truth, up to well-defined ambiguities. Within representation learning, identifiability has been studied mostly in the context of independent component analysis (ICA) , which assumes that the observed data results from mixing unobserved independent random variables referred to as sources. The aim is to recover the sources based on the observed mixtures alone, also termed blind source separation (BSS). A major obstacle to BSS is that, in the nonlinear case, independent component estimation does not necessarily correspond to recovering the true sources: it is possible to give counterexamples where the observations are transformed into components which are independent, yet still mixed with respect to the true sources . In other words, nonlinear ICA is not identifiable.
In order to achieve identifiability, a growing body of research postulates additional supervision or structure in the data generating process, often in the form of auxiliary variables . In the present work, we investigate a different route to identifiability by drawing inspiration from the field of causal inference which has provided useful insights for a number of machine learning tasks, including semi-supervised , transfer , reinforcement , and unsupervised learning. To this end, we interpret the ICA mixing as a causal process and apply the principle of independent causal mechanisms (ICM) which postulates that the generative process consists of independent modules which do not share information . In this context, “independent” does not refer to statistical independence of random variables, but rather to the notion that the distributions and functions composing the generative process are chosen independently by Nature . While a formalisation of ICM in terms of algorithmic (Kolmogorov) complexity exists, it is not computable, and hence applying ICM in practice requires assessing such non-statistical independence with suitable domain specific criteria . The goal of our work is thus to constrain the nonlinear ICA problem, in particular the mixing function, via suitable ICM measures, thereby ruling out common counterexamples to identifiability which intuitively violate the ICM principle.
Motivating example. To build intuition, we turn to a famous example of ICA and BSS: the cocktail party problem, illustrated in Fig. 1 (Left). Here, a number of conversations are happening in parallel, and the task is to recover the individual voices from the recorded mixtures . The mixing or recording process is primarily determined by the room acoustics and the locations at which microphones are placed. Moreover, each speaker influences the recording through their positioning in the room, and we may think of this influence as . Our independence postulate then amounts to stating that the speakers’ positions are not fine-tuned to the room acoustics and microphone placement, or to each other, i.e., the contributions should be independent (in a non-statistical sense).For additional intuition and possible violations in the context of the cocktail party problem, see § B.4.
Our approach. We formalise this notion of independence between the contributions of each source to the mixing process (i.e., the columns of the Jacobian matrix of partial derivatives) as an orthogonality condition, see Fig. 1 (Right). Specifically, the absolute value of the determinant , which describes the local change in infinitesimal volume induced by mixing the sources, should factorise or decompose as the product of the norms of its columns. This can be seen as a decoupling of the local influence of each partial derivative in the pushforward operation (mixing function) mapping the source distribution to the observed one, and gives rise to a novel framework which we term independent mechanism analysis (IMA). IMA can be understood as a refinement of the ICM principle that applies the idea of independence of mechanisms at the level of the mixing function.
Contributions. The structure and contributions of this paper can be summarised as follows:
we review well-known obstacles to identifiability of nonlinear ICA (§ 2.1), as well as existing ICM criteria (§ 2.2), and show that the latter do not sufficiently constrain nonlinear ICA (§ 3);
we propose a more suitable ICM criterion for unsupervised representation learning which gives rise to a new framework that we term independent mechanism analysis (IMA) (§ 4); we provide geometric and information-theoretic interpretations of IMA (§ 4.1), introduce an IMA contrast function which is invariant to the inherent ambiguities of nonlinear ICA (§ 4.2), and show that it rules out a large class of counterexamples and is consistent with existing identifiability results (§ 4.3);
we experimentally validate our theoretical claims and propose a regularised maximum-likelihood learning approach based on the IMA constrast which outperforms the unregularised baseline (§ 5); additionally, we introduce a method to learn nonlinear ICA solutions with triangular Jacobian and a metric to assess BSS which can be of independent interest for the nonlinear ICA community.
Background and preliminaries
Our work builds on and connects related literature from the fields of independent component analysis (§ 2.1) and causal inference (§ 2.2). We review the most important concepts below.
Assume the following data-generating process for independent component analysis (ICA)
If the true model belongs to the model class , then -identifiability ensures that any model in learnt from (infinite amounts of) data will be -equivalent to the true one. An example is linear ICA which is identifiable up to permutation and rescaling of the sources on the subspace of pairs of (i) invertible matrices (constraint on ) and (ii) factorizing densities for which at most one is Gaussian (constraint on ) , see Appendix A for a more detailed account.
In the nonlinear case (i.e., without constraints on ), identifiability is much more challenging. If and are independent, then so are and for any functions and . In addition to permutation-ambiguity, such element-wise can therefore not be resolved either. We thus define the desired form of identifiability for nonlinear BSS as follows.
The equivalence relation on defined as in Defn. 2.1 is given by
where is a permutation and is an invertible, element-wise function.
A fundamental obstacle—and a crucial difference to the linear problem—is that in the nonlinear case, different mixtures of and can be independent, i.e., solving ICA is not equivalent to solving BSS. A prominent example of this is given by the Darmois construction .
Let be an orthogonal matrix, and denote by and the element-wise CDFs of a smooth, factorised density and of a Gaussian, respectively. Then the “rotated-Gaussian” MPA is
2 Causal inference and the principle of independent causal mechanisms (ICM)
Rather than relying only on additional assumptions on (e.g., via auxiliary variables), we seek to further constrain (1) by also placing assumptions on the set of mixing functions . To this end, we draw inspiration from the field of causal inference . Of central importance to our approach is the Principle of Independent Causal Mechanisms (ICM) .
The causal generative process of a system’s variables is composed of autonomous modules that do not inform or influence each other.
These “modules” are typically thought of as the conditional distributions of each variable given its direct causes. Intuitively, the principle then states that these causal conditionals correspond to independent mechanisms of nature which do not share information. Crucially, here “independent” does not refer to statistical independence of random variables, but rather to independence of the underlying distributions as algorithmic objects. For a bivariate system comprising a cause and an effect , this idea reduces to an independence of cause and mechanism, see Fig. 2(c). One way to formalise ICM uses Kolmogorov complexity as a measure of algorithmic information .
However, since Kolmogorov complexity is is not computable, using ICM in practice requires assessing Principle 2.6 with other suitable proxy criteria .“This can be seen as an algorithmic analog of replacing the empirically undecidable question of statistical independence with practical independence tests that are based on assumptions on the underlying distribution” . Allowing for deterministic relations between cause (sources) and effect (observations), the criterion which is most closely related to the ICA setting in (1) is information-geometric causal inference (IGCI) .For a similar criterion which assumes linearity and its relation to linear ICA, see § B.1. IGCI assumes a nonlinear relation and formulates a notion of independence between the cause distribution and the deterministic mechanism (which we think of as a degenerate conditional ) via the following condition (in practice, assumed to hold approximately),
where is the Jacobian matrix and the absolute value of the determinant. can be understood as the covariance between and (viewed as r.v.s on the unit cube w.r.t. the Lebesgue measure), so that rules out a form of fine-tuning between and . As its name suggests, IGCI can, from an information-geometric perspective, also be seen as an orthogonality condition between cause and mechanism in the space of probability distributions , see § B.2, particularly eq. 19 for further details.
Existing ICM measures are insufficient for nonlinear ICA
Our aim is to use the ICM Principle 2.6 to further constrain the space of models and rule out common counterexamples to identifiability such as those presented in § 2.1. Intuitively, both the Darmois construction (4) and the rotated Gaussian MPA (5) give rise to “non-generic” solutions which should violate ICM: the former, , due the triangular Jacobian of (see Remark 2.4), meaning that each observation only depends on a subset of the inferred independent components , and the latter, , due to the dependence of on (5).
However, the ICM criteria described in § 2.2 were developed for the task of cause-effect inference where both variables are observed. In contrast, in this work, we consider an unsupervised representation learning task where only the effects (mixtures ) are observed, but the causes (sources ) are not. It turns out that this renders existing ICM criteria insufficient for BSS: they can easily be satisfied by spurious solutions which are not equivalent to the true one. We can show this for IGCI. Denote by the class of nonlinear ICA models satisfying IGCI (6). Then the following negative result holds.
(1) is not -identifiable on .
IGCI (6) is satisfied when is uniform. However, the Darmois construction (4) yields uniform sources, see Fig. 2(a). This means that , so IGCI can be satisfied by solutions which do not separate the sources in the sense of Defn. 2.2, see footnote 2 and .∎
As illustrated in Fig. 2(c), condition (6) and other similar criteria enforce a notion of “genericity” or “decoupling” of the mechanism w.r.t. the observed input distribution.In fact, many ICM criteria can be phrased as special cases of a unifying group-invariance framework . They thus rely on the fact that the cause (source) distribution is informative, and are generally not invariant to reparametrisation of the cause variables. In the (nonlinear) ICA setting, on the other hand, the learnt source distribution may be fairly uninformative. This poses a challenge for existing ICM criteria since any mechanism is generic w.r.t. an uninformative (uniform) input distribution.
Independent mechanism analysis (IMA)
As argued in § 3, enforcing independence between the input distribution and the mechanism (Fig. 2(c)), as existing ICM criteria do, is insufficient for ruling out spurious solutions to nonlinear ICA. We therefore propose a new ICM-inspired framework which is more suitable for BSS and which we term independent mechanism analysis (IMA).The title of the present work is thus a reverence to Pierre Comon’s seminal 1994 paper . All proofs are provided in Appendix C.
As motivated using the cocktail party example in § 1 and Fig. 1 (Left), our main idea is to enforce a notion of independence between the contributions or influences of the different sources on the observations as illustrated in Fig. 2(d)—as opposed to between the source distribution and mixing function, cf. Fig. 2(c). These contributions or influences are captured by the vectors of partial derivatives . IMA can thus be understood as a refinement of ICM at the level of the mixing : in addition to statistically independent components , we look for a mixing with contributions which are independent, in a non-statistical sense which we formalise as follows.
The mechanisms by which each source influences the observed distribution, as captured by the partial derivatives , are independent of each other in the sense that for all :
Geometric interpretation. Geometrically, the IMA principle can be understood as an orthogonality condition, as illustrated for in Fig. 1 (Right). First, the vectors of partial derivatives , for which the IMA principle postulates independence, are the columns of . thus measures the volume of the dimensional parallelepiped spanned by these columns, as shown on the right. The product of their norms, on the other hand, corresponds to the volume of an -dimensional box, or rectangular parallelepiped with side lengths , as shown on the left. The two volumes are equal if and only if all columns of are orthogonal. Note that (7) is trivially satisfied for , i.e., if there is no mixing, further highlighting its difference from ICM for causal discovery.
Information-geometric interpretation and comparison to IGCI. The additive contribution of the sources’ influences in (7) suggests their local decoupling at the level of the mechanism . Note that IGCI (6), on the other hand, postulates a different type of decoupling: one between and . There, dependence between cause and mechanism can be conceived as a fine tuning between the derivative of the mechanism and the input density. The IMA principle leads to a complementary, non-statistical measure of independence between the influences of the individual sources on the vector of observations. Both the IGCI and IMA postulates have an information-geometric interpretation related to the influence of (“non-statistically”) independent modules on the observations: both lead to an additive decomposition of a KL-divergence between the effect distribution and a reference distribution. For IGCI, independent modules correspond to the cause distribution and the mechanism mapping the cause to the effect (see (19) in § B.2). For IMA, on the other hand, these are the influences of each source component on the observations in an interventional setting (under soft interventions on individual sources), as measured by the KL-divergences between the original and intervened distributions. See § B.3, and especially (22), for a more detailed account.
We finally remark that while recent work based on the ICM principle has mostly used the term “mechanism” to refer to causal Markov kernels or structural equations , we employ it in line with the broader use of this concept in the philosophical literature.See Table 1 in for a long list of definitions from the literature. To highlight just two examples, states that “Causal processes, causal interactions, and causal laws provide the mechanisms by which the world works; to understand why certain things happen, we need to see how they are produced by these mechanisms”; and states that “Mechanisms are events that alter relations among some specified set of elements”. Following this perspective, we argue that a causal mechanism can more generally denote any process that describes the way in which causes influence their effects: the partial derivative thus reflects a causal mechanism in the sense that it describes the infinitesimal changes in the observations , when an infinitesimal perturbation is applied to .
2 Definition and useful properties of the IMA contrast
We now introduce a contrast function based on the IMA principle (7) and show that it possesses several desirable properties in the context of nonlinear ICA. First, we define a local contrast as the difference between the two integrands of (7) for a particular value of the sources .
The local IMA contrast of at a point is given by
This corresponds to the left KL measure of diagonality for .
The local IMA contrast quantifies the extent to which the IMA principle is violated at a given point . We summarise some of its properties in the following proposition.
The local IMA contrast defined in (8) satisfies:
, with equality if and only if all columns of are orthogonal.
is invariant to left multiplication of by an orthogonal matrix and to right multiplication by permutation and diagonal matrices.
Property (i) formalises the geometric interpretation of IMA as an orthogonality condition on the columns of the Jacobian from § 4.1, and property (ii) intuitively states that changes of orthonormal basis and permutations or rescalings of the columns of do not affect their orthogonality. Next, we define a global IMA contrast w.r.t. a source distribution as the expected local IMA contrast.
The global IMA contrast of w.r.t. is given by
The global IMA contrast thus quantifies the extent to which the IMA principle is violated for a particular solution to the nonlinear ICA problem. We summarise its properties as follows.
The global IMA contrast from (9) satisfies:
Property (i) is the distribution-level analogue to (i) of Prop. 4.4 and only allows for orthogonality violations on sets of measure zero w.r.t. . This means that can only be zero if is an orthogonal coordinate transformation almost everywhere , see Fig. 3 for an example. We particularly stress property (ii), as it precisely matches the inherent indeterminacy of nonlinear ICA: is blind to reparametrisation of the sources by permutation and element wise transformation.
We now show that, under suitable assumptions on the generative model (1), a large class of spurious solutions—such as those based on the Darmois construction (4) or measure preserving automorphisms such as from (5) as described in § 2.1—exhibit nonzero IMA contrast. Denote the class of nonlinear ICA models satisfying (7) (IMA) by . Our first main theoretical result is that, under mild assumptions on the observations, Darmois solutions will have strictly positive , making them distinguishable from those in .
Assume the data generating process in (1) and assume that for some . Then any Darmois solution based on as defined in (4) satisfies . Thus a solution satisfying can be distinguished from based on the contrast .
The proof is based on the fact that the Jacobian of is triangular (see Remark 2.4) and on the specific form of (4). A specific example of a mixing process satisfying the IMA assumption is the case where is a conformal (angle-preserving) map.
Under assumptions of Thm. 4.7, if additionally is a conformal map, then for any due to Prop. 4.6 (i), see Defn. 4.8. Based on Thm. 4.7, is thus distinguishable from Darmois solutions .
This is consistent with a result that proves identifiability of conformal maps for and conjectures it in general .Note that Corollary 4.9 holds for any dimensionality . However, conformal maps are only a small subset of all maps for which , as is apparent from the more flexible condition of Prop. 4.6 (i), compared to the stricter Defn. 4.8.
Consider the non-conformal transformation from polar to Cartesian coordinates (see Fig. 3), defined as with independent sources , with and .For different , can be made to have independent Gaussian components (, II.B), and -identifiability is lost; this shows that the assumption of Thm. 4.7 that for some is crucial. Then, and for any Darmois solution —see Appendix D for details.
Finally, for the case in which the true mixing is linear, we obtain the following result.
In a “blind” setting, we may not know a priori whether the true mixing is linear or not, and thus choose to learn a nonlinear unmixing. Corollary 4.11 shows that, in this case, Darmois solutions are still distinguishable from the true mixing via . Note that unlike in Corollary 4.9, the assumption that for some is not required for Corollary 4.11. In fact, due to Theorem 11 of , it follows from the assumed linear ICA model with non-Gaussian sources, and the fact that the mixing matrix is not the product of a diagonal and a permutation matrix (see also Appendix A).
Having shown that the IMA principle allows to distinguish a class of models (including, but not limited to conformal maps) from Darmois solutions, we next turn to a second well-known counterexample to identifiability: the “rotated-Gaussian” MPA (5) from Defn. 2.5. Our second main theoretical result is that, under suitable assumptions, this class of MPAs can also be ruled out for “non-trivial” .
Let and assume that is a conformal map. Given , assume additionally that at least one non-Gaussian whose associated canonical basis vector is not transformed by into another canonical basis vector . Then .
Thm. 4.12 states that for conformal maps, applying the transformation at the level of the sources leads to an increase in except for very specific rotations that are “fine-tuned” to in the sense that they permute all non-Gaussian sources with another . Interestingly, as for the linear case, non-Gaussianity again plays an important role in the proof of Thm. 4.12.
Experiments
Our theoretical results from § 4 suggest that is a promising contrast function for nonlinear blind source separation. We test this empirically by evaluating the of spurious nonlinear ICA solutions (§ 5.1), and using it as a learning objective to recover the true solution (§ 5.2).
We sample the ground truth sources from a uniform distribution in ; the reconstructed sources are also mapped to the uniform hypercube as a reference measure via the CDF transform. Unless otherwise specified, the ground truth mixing is a Möbius transformation (i.e., a conformal map) with randomly sampled parameters, thereby satisfying Principle 4.1. In all of our experiments, we use JAX and Distrax . For additional technical details, equations and plots see Appendix E. The code to reproduce our experiments is available at this link.
Learning the Darmois construction. To learn the Darmois construction from data, we use normalising flows, see . Since Darmois solutions have triangular Jacobian (Remark 2.4), we use an architecture based on residual flows which we constrain such that the Jacobian of the full model is triangular. This yields an expressive model which we train effectively via maximum likelihood.
of Darmois solutions. To check whether Darmois solutions (learnt from finite data) can be distinguished from the true one, as predicted by Thm. 4.7, we generate random mixing functions for , compute the values of learnt solutions, and find that all values are indeed significantly larger than zero, see Fig. 4 (a). The same holds for higher dimensions, see Fig. 4 (b) for results with random mixings for : with higher dimensionality, both the mean and variance of the distribution for the learnt Darmois solutions generally attain higher values.the latter possibly due to the increased difficulty of the learning task for larger We confirmed these findings for mappings which are not conformal, while still satisfying (7), in § E.5.
of MPAs. We also investigate the effect on of applying an MPA from (5) to the true solution or a learnt Darmois solution. Results for dim. for different rotation matrices (parametrised by the angle ) are shown in Fig. 4 (c). As expected, the behavior is periodic in , and vanishes for the true solution (blue) at multiples of , i.e., when is a permutation matrix, as predicted by Thm. 4.12. For the learnt Darmois solution (red, dashed) remains larger than zero.
values for random MLPs. Lastly, we study the behavior of spurious solutions based on the Darmois construction under deviations from our assumption of for the true mixing function. To this end, we use invertible MLPs with orthogonal weight initalisation and leaky_tanh activations as mixing functions; the more layers are added to the mixing MLP, the larger a deviation from our assumptions is expected. We compare the true mixing and learnt Darmois solutions over realisations for each , . Results are shown in figure Fig. 4 (d): the of the mixing MLPs grows with ; still, the one of the Darmois solution is typically higher.
Summary. We verify that spurious solutions can be distinguished from the true one based on .
Results. In Fig. 4 (Top), we show an example of the distortion induced by different spurious solutions for , and contrast it with a solution learnt using our proposed objective (rightmost plot). Visually, we find that the -regularised solution (with ) recovers the true sources most faithfully. Quantitative results for 50 learnt models for each and are summarised in Fig. 5 (see Appendix E for additional plots) . As indicated by the KL divergence values (left), most trained models achieve a good fit to the data across all values of .models with have high outlier KL values, seemingly less pronounced for nonzero values of We observe that using (i.e., ) is beneficial for BSS, both in terms of our nonlinear Amari distance (center, lower is better) and MCC (right, higher is better), though we do not observe a substantial difference between and .In § E.5, we also show that our method is superior to a linear ICA baseline, FastICA .
Summary: can be a useful learning signal to recover the true solution.
Discussion
Assumptions on the mixing function. Instead of relying on weak supervision in the form of auxiliary variables , our IMA approach places additional constraints on the functional form of the mixing process. In a similar vein, the minimal nonlinear distortion principle proposes to favor solutions that are as close to linear as possible. Another example is the post-nonlinear model , which assumes an element-wise nonlinearity applied after a linear mixing. IMA is different in that it still allows for strongly nonlinear mixings (see, e.g., Fig. 3) provided that the columns of their Jacobians are (close to) orthogonal. In the related field of disentanglement , a line of work that focuses on image generation with adversarial networks similarly proposes to constrain the “generator” function via regularisation of its Jacobian or Hessian , though mostly from an empirically-driven, rather than from an identifiability perspective as in the present work.
Towards identifiability with . The IMA principle rules out a large class of spurious solutions to nonlinear ICA. While we do not present a full identifiability result, our experiments show that can be used to recover the BSS equivalence class, suggesting that identifiability might indeed hold, possibly under additional assumptions—e.g., for conformal maps .
IMA and independence of cause and mechanism. While inspired by measures of independence of cause and mechanism as traditionally used for cause-effect inference , we view the IMA principle as addressing a different question, in the sense that they evaluate independence between different elements of the causal model. Any nonlinear ICA solution that satisfies the IMA Principle 4.1 can be turned into one with uniform reconstructed sources—thus satisfying IGCI as argued in § 3—through composition with an element-wise transformation which, according to Prop. 4.6 (ii), leaves the value unchanged. Both IGCI (6) and IMA (7) can therefore be fulfilled simultaneosly, while the former on its own is inconsequential for BSS as shown in Prop. 3.1.
BSS through algorithmic information. Algorithmic information theory has previously been proposed as a unifying framework for identifiable approaches to linear BSS , in the sense that commonly-used contrast functions could, under suitable assumptions, be interpreted as proxies for the total complexity of the mixing and the reconstructed sources. However, to the best of our knowledge, the problem of specifying suitable proxies for the complexity of nonlinear mixing functions has not yet been addressed. We conjecture that our framework could be linked to this view, based on the additional assumption of algorithmic independence of causal mechanisms , thus potentially representing an approach to nonlinear BSS by minimisation of algorithmic complexity.
ICA for causal inference & causality for ICA. Past advances in ICA have inspired novel causal discovery methods . The present work constitutes, to the best of our knowledge, the first effort to use ideas from causality (specifically ICM) for BSS. An application of the IMA principle to causal discovery or causal representation learning is an interesting direction for future work.
Conclusion. We introduce IMA, a path to nonlinear BSS inspired by concepts from causality. We postulate that the influences of different sources on the observed distribution should be approximately independent, and formalise this as an orthogonality condition on the columns of the Jacobian. We prove that this constraint is generally violated by well-known spurious nonlinear ICA solutions, and propose a regularised maximum likelihood approach which we empirically demonstrate to be effective in recovering the true solution. Our IMA principle holds exactly for orthogonal coordinate transformations, and is thus of potential interest for learning spatial representations , robot dynamics , or physics problems where orthogonal reference frames are common .
Acknowledgements
The authors thank Aapo Hyvärinen, Adrián Javaloy Bornás, Dominik Janzing, Giambattista Parascandolo, Giancarlo Fissore, Nasim Rahaman, Patrick Burauel, Patrik Reizinger, Paul Rubenstein, Shubhangi Ghosh, and the anonymous reviewers for helpful comments and discussions.
Funding Transparency Statement
This work was supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039B; and by the Machine Learning Cluster of Excellence, EXC number 2064/1 - Project number 390727645.
References
Checklist
Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]
Did you describe the limitations of your work? [Yes] See § 6, where we discuss limitations of our theory (e.g. lines 354-357) and open questions.
Did you discuss any potential negative societal impacts of your work? [N/A] Our work is mainly theoretical, and we believe it does not bear immediate negative societal impacts.
Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]
If you are including theoretical results…
Did you state the full set of assumptions of all theoretical results? [Yes] We formally define the problem setting in § 2 and § 3, and transparently state and discuss our assumptions in § 4.
Did you include complete proofs of all theoretical results? [Yes] Due to space constraints, full proofs and detailed explanations are mainly reported in appendix C and appendix D; the proof of Prop. 3.1 is given in the main text.
Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] The code, data, and the configuration files are included in the supplemental material.
Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] All details, including hyperparameters, seed for random number generators, etc. are specified in the configuration files which will be included in the supplemental.
Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [Yes] We visualized the distribution of the considered quantities via histograms and violin plots, see e.g. Figure 4 and Figure 5.
Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] They are specified appendix E.
If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…
If your work uses existing assets, did you cite the creators? [Yes] We use the Python libraries JAX and Distrax and cited the creators in the article.
Did you mention the license of the assets? [Yes] Both packages have Apache License 2.0; we report this in the appendices.
Did you include any new assets either in the supplemental material or as a URL? [Yes] Implementations of our proposed methods and metrics will be provided in the supplemental.
Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A]
Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A]
If you used crowdsourcing or conducted research with human subjects…
Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]
Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]
Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]
Overview
Appendix A contains further elaboration on the notion of identifiability as used in the present work, as well as connections to linear ICA.
Appendix B contains additional discussion of existing ICM criteria and their relation to IMA.
Appendix C presents the full proofs for all theoretical results from the main paper.
Appendix D contains a worked out computation of the value of for the mapping from radial to Cartesian coordinates.
Appendix E contains experimental details and additional results.
Appendix F contains additional background on conformal maps and Möbius transformations
Appendix A Additional background on identifiability and linear ICA
In this Appendix, we provide additional background on the notion of identifiability and illustrate it using the example of linear ICA.
Traditionally, identifiability for a class of models for observed data parametrised by is expressed as the condition that there needs to be a one-to-one mapping between the space of models and the space of parameters, i.e., the model class is said to be identifiable if
However, the equality on the RHS of (10) is a very strong condition which makes this type of (strong or unique) identifiability impractical for many settings. For example, in the case of (linear or nonlinear) ICA, the ordering of the sources cannot be determined, so strong identifiability in the sense of (10) is infeasible.
The equality in parameter space on the RHS of the implication in (10) is therefore sometimes replaced by an equivalence relation , as is also the case for our Defn. 2.1. An equivalence relation on a set is a binary relation between pairs of elements of which satisfies the following three properties:
Reflexivity: , .
Symmetry: , .
Transitivity: .
An equivalence relation on a set imposes a partition into disjoint subsets. Each such subset corresponds to an equivalence class, i.e., the collection of all elements which are -related to each other; for example, denotes the equivalence class containing the element .
A trivial example of an equivalence relation is equality (). More useful examples in the context of ICA are equivalence up to permutation, rescaling, or scalar transformation.
Defining an appropriate equivalence class for the problem at hand therefore allows us to specify exactly the type of indeterminancies which cannot be resolved and up to which the true generative process can be recovered. As argued in § 2, for nonlinear ICA, the desired notion of identifiability—in the sense of the strongest feasible type of identifiable that is possible without further (parametric) assumptions—is captured by from Defn. 2.2. We give another example for linear ICA in § A.2.
Since the generative process of nonlinear ICA (1) is determined by the choice of mixing function and source distribution, the space from (10), in this case, corresponds to the product space of the space of mixing functions and source distributions . Moreover, the pushforward density in Defn. 2.1 corresponds to the density of the observed mixtures , or in (10).
We deliberately choose to define identifiability and to express the observed distribution in terms of the source distribution and the mixing function—as opposed to in terms of the observed distribution and the unmixing function as in some prior work —because this is aligned with the causal direction of data generation, and thus more consistent with the causal perspective at nonlinear BSS taken in the present work. We also believe that, in this framework, separate constraints on the space of mixing functions and source distributions are expressed more naturally.
Next, we illustrate the above ideas for the well-studied case of linear ICA.
A.2 Identifiability of linear ICA
Linear ICA corresponds to the setting in which a linear mixing is applied to independent sources, i.e.,
Additionally, we can assume w.l.o.g. that the mixing matrix is orthogonal (), because we can always whiten first through an invertible linear transformation and obtain an orthogonal mixing , as explained in more detail in § A.3.
Now suppose that the reconstructed sources
Thm. A.1 shows that the two ambiguities deemed unresolvable (scale and ordering of the sources) are, in fact, the only ambiguities, as long as at most one of the is Gaussian. That is, linear ICA is identifiable up to rescaling and permutation of the sources, i.e., linearly transforming the observations into independent components is equivalent to separating the sources.
More formally, in terms of an equivalence relation, if we take from (2) as the space of invertible matrices and as the space of source distributions with at most one Gaussian marginal, then linear ICA is -identifiable on where the equivalence relation on is defined as
Two other deviations from a Gaussian i.i.d. setting lead to identifiability: nonstationarity and time correlation . A general information-geometric framework links these three different routes to identifiability .
A.3 Whitening in the context of linear ICA
For completeness, we give a brief account of the role of whitening in linear ICA, which was mentioned in A.2 and which again plays a role in B.1. The following exposition is partly based on , §7.4.2.
A zero-mean random vector, say , is said to be white if its components are uncorrelated and their variances equal unity. In other words, the covariance matrix of is equal to the identity matrix:
It is always possible to whiten a zero-mean random vector through a linear operation,
As an example, a popular method for whitening uses the eigenvalue decomposition (EVD) of the covariance matrix,
Whitening is only half ICA. Assume a linear ICA model,
Note that whitening does not solve linear ICA, since uncorrelatedness is weaker than independence. To see this, consider any orthogonal transformation of :
Due to the orthogonality of we have
so, is white as well. Thus, we cannot tell if the independent components are given by or using the whiteness property alone. Since could be any orthogonal transformation of whitening gives the independent components only up to an orthogonal transformation.
Whitening thus “solves half of the problem of ICA”. Because whitening is a very simple and standard procedure—much simpler than any ICA algorithm—it is a good idea to reduce the complexity of the problem this way. The remaining half of the parameters has to be estimated by some other method.
Appendix B Existing ICM criteria and their relationship to ICA and IMA
We now provide additional discussion of the ICM principle and its connection to ICA and IMA. First, we introduce a linear ICM criterion and discuss its relation with linear ICA in § B.1.
As mentioned in § 2.2, besides IGCI, another existing ICM criterion that is closely related to ICA due to also assuming a deterministic relation between cause and effect is the trace method . The trace method assumes a linear relationship,
and formulates ICM as an “independence” between the covariance matrix of and the mechanism (which, as for IGCI, we can again think of as a degenerate conditional ) via the condition
where denotes the renormalized trace. Intuitively, this condition (17) rules out a fine-tuning of to the eigenvectors of which would violate the assumption of no shared information between the cause distribution (specifically, its covariance structure) and the mechanism.
As with IGCI and nonlinear ICA, it can be seen by comparing (16) and (11) that the trace method assumes the same generative model as linear ICA (where the cause corresponds to the independent sources and the effect to the observed mixtures ). While the focus of the present work is on nonlinear ICA, we briefly discuss the usefulness of the trace method as a constraint for achieving identifiability in a linear ICA setting.
As is clear from (17), the trace condition is trivially satisfied if the covariance matrix of the sources (causes) is the identity, . However, as explained in § A.3, in the context of linear ICA this can easily be achieved by whitening the data. As with IGCI, the trace method was developed for cause-effect inference where both variables are observed, and thus relies on the observed cause distribution being informative. This renders is unsuitable (on its own) to constrain the unsupervised representation learning problem of linear ICA problem where the sources are unobserved.
Note, however, that this is qualitatively different from the IGCI argument presented in § 3, as whitening on its own does not necessarily lead to independent variables, but only uncorrelated ones, and thus does not solve linear ICA—unlike the Darmois construction in the case of nonlinear ICA which also yields independent components.
B.2 Information geometric interpretation of the ICM principle
There is a well-established connection between IGCI and the trace method . At the heart of this derivation lies an information-geometric interpretation of the ICM principle for probability distributions, which we sketch in this section. First, we need to review some basic concepts.
An interesting property of the KL divergence is its invariance to reparametrisation. Consider an invertible transformation , mapping random variables and to and , respectively (the domains and codomains being arbitrary spaces, e.g., discrete or Euclidean of arbitrary dimension). Then the KL divergence between and is preserved by the pushforward operation implemented by , such that
Consider a deterministic causal relationship of the form , and denote by and the marginal distributions of the cause and the effect , respectively. The “irregularity” of each distribution can be quantified by evaluating their divergence to a reference set of “regular” distributions,Here “regular” is only meant in an intuitive sense, not implying any further mathematical notion. If is the set of Gaussians, for instance, the distance from measures non-Gaussianity.
Let us assume that these infima are reached at a unique point, their projections onto :
As elaborated in [46, §4], the choice of is context-dependent. For example, in the context of the trace method , and are assumed to be -dimensional multivariate Gaussian random vectors, and is taken as the set of multivariate isotropic Gaussian distributions. In contrast, when IGCI is applied in contexts where the considered mechanism is a deterministic non-linear diffeomorphism, the reference distributions are typically uniform distributions .
Overall, it can be shown that the independence postulate underlying these approaches leads to the following decomposition of the irregularity of (see [46, Thm. 2]):
where denotes the distribution of , i.e., the hypothetical distribution of the effect that would be obtained if the cause were replaced by the random variable (which corresponds to the closest regularly distributed random variable to ).
Since applying the bijection preserves the KL divergences, see (18), we can obtain the equivalent relation
This relation can be interpreted as an orthogonality principle in information space by considering the KL divergences as a generalization of the squared Euclidean norm for the difference vectors , and . It can thus be viewed as a Pythagorean theorem in the space of distributions, see Fig. 6 for an illustration.
The orthogonality principle (19) thus captures a decomposition of the irregularity of on the LHS into the sum of two irregularities on the RHS: the irregularity of , and the term which measures the irregularity of the mechanism indirectly, via the “irregularity” of the distribution resulting from applying to a regular distribution .
Overall, the decomposition (19) links the postulate of independence between the cause distribution, on the one hand, and the mechanism, on the other hand, to an orthogonality of their irregularities in information space (namely the statistical manifold of information geometry). As proposed in , this can be intuitively interpreted as a geometric form of independence if we assume that Nature chooses such irregularities independently of each other, and “isotropically” in a high-dimensional subspace of irregularities.
While, to date, we are not aware of similar results in the context of information geometry (i.e., on the statistical manifold), this intuition is supported by concentration of measure results in Euclidean spaces. Indeed, in high-dimensions, it is likely that two vectors are close to orthogonal if they are chosen independently according to a uniform prior .
We will take inspiration of the decomposition (19) to justify IMA in the following section.
B.3 Decoupling of the influences in IMA and comparison with IGCI
In contrast to § B.2, in this section we will, for notational consistency with the main paper, assume that all distributions have a density with respect to the Lebesgue measure, and thus consider, with a slight abuse of notation, that the KL divergence is a distance between two densities on the relevant support, such that
In line with the information-geometric interpretation of IGCI presented in § B.2, we also consider an interpretation of IMA in information space. We consider the KL-divergence between the observed density of and an interventional distribution of , resulting from a soft intervention that replaces the mixing function with another mixing . We take as a measure of the causal effect of the soft intervention (or perturbation) that turns into —similarly to how is used as a measure of the irregularity of the effect distribution in the context of IGCI (§ B.2).
As we will show, under suitable assumptions, the functional form imposed on by the IMA Principle 4.1 can lead to a decomposition of the causal effect of an intervention on the mechanism into a sum of terms, corresponding to the causal effects of separate soft interventions on the mechanisms associated to each source. In contrast, IGCI decomposes irregularities of the effect distribution into two terms, one irregularity of the cause and one irregularity of the mechanism.
Assume satisfies the IMA principle. We consider interventions performed through the element-wise transformation such that
This can be seen as a composition of soft interventions on each individual source component , implemented through univariate smooth diffeomorphisms , such that
and (in arbitrary order, since the individual commute). This soft intervention can be seen as turning the random variable into , yielding the intervened observations . Alternatively, the intervention on can be implemented by replacing by —i.e., . Notably, since satisfies the IMA principle, so does (due to Prop. 4.6, (ii), since is an element-wise nonlinearity). Moreover, the partial derivatives of the intervened function are given by
The classical change of variable formula for bijection yields the expression of the pushforward density of as
Let us now compute the KL divergence between the intervened and observed distribution,
Expressing the density of the observed variables as a pushforward of the density of the sources, and without additional assumptions on and besides smoothness and invertibility, we get,
We now consider a factorization of over a directed acyclic graph (DAG), such that
where denotes the components associated to the parents of node in the DAG. Because is an element-wise transformation the factorization will be the same for .
If we now additionally assume that and satisfy the IMA postulate, we get
By reparameterizing the integral in terms of the source coordinates, we get (using )
such that the divergence can be written as a sum of terms, each associated to the intervention on a mechanism . Positivity of these terms would suggest that we can interpret each of them as quantifying the individual contribution of a soft intervention applied to the original sources.
In the following, we propose a justification for the positivity of these terms in a restricted setting where only the leaf nodes of the graph are intervened on (with ).A leaf node in a DAG is one that does not have any descendants. In the special case of independent sources, all nodes are leaves and .
Under this assumption, we consider (without loss of generality) an ordering of the nodes such that the first nodes are the leaf nodes in the DAG. Then we argue that the terms of the right-hand side of (21) associated to leaf nodes () are positive, as they correspond to the expectations of KL-divergences. Indeed, taking one of the first terms, denoted , we have the factorization
where does not depend on because node is a leaf node. Moreover, as non-leaf nodes are not intervened on, the transformation does not modify the value of any parent variables in these factorizations. As a consequence, the integral can be computed as an iterated integral with respect to and , where denotes the vector including all source variables but , such that
Similarly, we get the expression of the pushfoward distribution from to the curve (using again the fact that parent variables are not intervened on, and thus left unchanged by )
These terms appear when rewriting the -th term (for a leaf variable) in (21) as a curvilinear integral:
The inner integral term can thus be interpreted as the KL divergence between two pushforward measures defined on by and , that we can denote by
To conclude, this implies that the causal effect of the soft intervention can be decomposed as the following sum of positive terms associated to interventions on each leaf variable, plus an additional term for the remaining non-leaf variables, which further simplifies (in comparison to (21)) due to the assumption that those variables are unintervened.
This expression suggests that the KL-divergences appearing in the first terms each reflect the causal effect of an intervention on the mechanism at the level of one single source coordinate , turning into . When the sources are jointly independent, we have and the right hand side of (22) contains only positive terms. An interesting direction for future work would be to analyse the remaining term in the case of non unconditionally independent sources.
In contrast to the decomposition (19) in the context of IGCI, the IMA decomposition (22) involves (expectations of) KL-divergence terms instead of two, each related to the intervention on the part of the mechanism that reflects the influence of a single source.
B.4 Independence of cause and mechanism and IMA
We now discuss an example in which a formalisation of the principle of independence of cause and mechanism is violated, and one in which the IMA principle is violated.
In the context of the Trace method , used in causal discovery, a technical example of fine-tuning can be constructed by taking a vector of i.i.d. random variables with arbitrary (not diagonal) covariance matrix as the cause, and by constructing the mechanism as a whitening matrix, turning the cause variables into uncorrelated (effect) variables. By doing so, the singular values and singular vectors of the matrix (the mechanism) are fine-tuned to the input covariance matrix (a property of the cause distribution), and such fine-tuning can be quantified via the Trace method (see , Section 1).
B.4.2 Violations of the IMA principle
As mentioned in § 3, an example of a mixing function which is non-generic according to the IMA principle is an autoregressive function, for example an autoregressive normalising flow , where the -th component of the observations only depends on the -th sources: intuitively, this would correspond to the unlikely cocktail-party setting where the -th microphone only picks up the voices of the first speakers. More precisely, as we show in Lemma C.1, this leads to positive value for such mixing.
A cocktail party (Fig. 1, left) may violate our IMA principle when the locations of several speakers and the room acoustics have been fine tuned to one another. This is for example the case in concert halls where the acoustics of the room have been fine-tuned to the position and configuration of multiple locations on the stage, where the sources (i.e., the voices of the actors or singers) are emitted—in order to make the listening experience as homogeneous as possible across the spectators (that is, the influence of each of the sources on the different listeners should not differ too much). This would lead to an increase in collinearity between the columns of the mixing’s Jacobian, thus violating the IMA principle.
Additionally, we recall that the ICM principle is often informally introduced by referencing the fine-tuning and non-generic viewpoints giving rise to certain visual illusions, such as the Beuchet chair (see , Section 2); in a similar vein, we can imagine that violations of the IMA principle in the cocktail party setting may be related to illusions in binaural hearing such as for example the Franssen effect, where the listener is tricked into incorrectly localizing a sound .
Appendix C Proofs
We now provide the proofs of all our theoretical results from the main paper.
Before giving the proof, it is useful to rewrite the local IMA constrast (8) as follows:
where the quantity in (23) is called the left KL measure of diagonality of the matrix (see Remark 4.3):
From (23), it can be seen that is a function of only through .
For ease of exposition, we denote the value of the Jacobian of evaluated at the point by . The two properties can then be proved as follows:
This is a consequence of Hadamard’s inequality, applied to the expression on the RHS of (8), which states that, for a matrix with columns , ; equality in Hadamard’s inequality is achieved iff. the vectors are orthogonal.
C.2 Proof of Prop. 4.6
From property (i) of Prop. 4.4, we know that . Hence, follows as a direct consequence of integrating the non-negative quantity .
Equality is attained iff. almost surely w.r.t. , which according to property (i) of Prop. 4.4 occurs iff. the columns of are orthogonal almost surely w.r.t. .
It remains to show that this is the case iff. can be written as , with and orthogonal and diagonal matrices, respectively. (To avoid confusion, note that orthogonal columns need not have unit norm, whereas an orthogonal matrix satisfies .)
The if is clear since right multiplication by a diagonal matrix merely re-scales the columns, and hence does not affect their orthogonality.
For the only if, let be any matrix with orthogonal columns , , and denote the column norms by . Further denote the normalised columns of by and let and be the orthogonal and diagonal matrices with columns and diagonal elements , respectively. Then .
where, for the second equality, we have used the fact that
since is an invertible tranformation (see, e.g., ). It thus suffices to show that
where we have repeatedly used the chain rule for Jacobians, as well as that ; that permutation is a linear operation, so for any ; and that (and thus ) is an element-wise transformation, so the Jacobian is a diagonal matrix .
The equality in (25) then follows from (26) by applying property (ii) of Prop. 4.4, according to which is invariant to right multiplication of the Jacobian by diagonal and permutation matrices.
Substituting (25) into the RHS of (24), we finally obtain
C.3 Remark on a similar condition to IMA, expressed in terms of the rows of the Jacobian
We remark that the condition imposed by the IMA Principle 4.1 needs to be expressed in terms of the columns of the Jacobian, and would not lead to a criterion with desirable properties for BSS if it were instead expressed in terms of its rows (which correspond to gradients of the ). One way to justify this is that, for the same condition expressed on the rows of the Jacobian, that is
property (ii) of Prop. 4.4 would not hold (because invariance would hold w.r.t. right, not left, multiplication with a diagonal matrix). As a consequence, the resulting global contrast would not be blind to reparametrisation of the source variables by permutation and element-wise invertible transformations, thereby not being a good contrast in the context of blind source separation.
C.4 Proof of Thm. 4.7
Before proving the main theorem, we first introduce some additional details on the Jacobian of the Darmois construction which will be important for the proof.
Consider the Darmois construction for ,
In the general case, the Jacobian of the Darmois construction will be
where the components of for all are defined by
It is additionally useful to introduce the following lemmas.
A function with triangular Jacobian has iff. its Jacobian is diagonal almost everywhere. Otherwise, .
Let have lower triangular Jacobian at , and denote . Then we have
where . Since the logarithm is a strictly monotonically increasing function and since
with equality iff. (i.e., iff. is a diagonal matrix), we must have iff. is diagonal.
is therefore equal to zero iff. has diagonal Jacobian almost everywhere, and it is strictly larger than zero otherwise. ∎
Let be a smooth function with diagonal Jacobian everywhere.
Consider the function for any . Suppose for a contradiction that depends on for some . Then there must be at least one point such that . However, this contradicts the assumption that is diagonal everywhere (since is an off-diagonal element for ). Hence, can only depend on for all , i.e., is an element wise function. ∎
First, the Jacobian of the Darmois construction is lower triangular , see (28).
Because CDFs are monotonic functions (strictly monotonically increasing given our assumptions on and ), is invertible.
We can thus apply the inverse function theorem (with ) to write
Since the inverse of a lower triangular matrix is lower triangular, we conclude that is lower triangular for all .
Now, according to Lemma C.1, we have , unless is diagonal almost everywhere.
Suppose for a contradiction that is diagonal almost everywhere.
Since and are smooth by assumption, so is the push-forward , and thus also (CDF of a smooth density) and its inverse . Hence, the partial derivatives , i.e., the elements of are continuous.
Consider an off-diagonal element for . Since these are zero almost everywhere, and because continuous functions which are zero almost everywhere must be zero everywhere, we conclude that everywhere for , i.e., the Jacobian is diagonal everywhere.
Hence, we conclude from Lemma C.2 that must be an element-wise function, .
Since has independent components by construction, it follows that and are independent for any .
However, this constitutes a contradiction to the assumption that x_{i}\not\mathchoice{\mathrel{\hbox to0.0pt{\displaystyle\perp\hss}\mkern 2.0mu{\displaystyle\perp}}}{\mathrel{\hbox to0.0pt{\textstyle\perp\hss}\mkern 2.0mu{\textstyle\perp}}}{\mathrel{\hbox to0.0pt{\scriptstyle\perp\hss}\mkern 2.0mu{\scriptstyle\perp}}}{\mathrel{\hbox to0.0pt{\scriptscriptstyle\perp\hss}\mkern 2.0mu{\scriptscriptstyle\perp}}}x_{j} for some .
We conclude that cannot be diagonal almost everywhere, and hence, by Lemma C.1, we must have . ∎
C.5 Proof of Corollary 4.9
C.6 Proof of Corollary 4.11
Since, by assumption, the mixing matrix is non-trivial (i.e., not the product of a diagonal and permutation matrix), and at most one of the is Gaussian, according to Thm. A.1 there must be at least one pair , with , such that .
We can then use the same argument as in the proof of Thm. 4.7 to show that the Darmois construction has nonzero , whereas the linear orthogonal transformation has orthogonal Jacobian, and thus by property (i) of Prop. 4.6. ∎
C.7 Proof of Thm. 4.12
For notational convenience, we denote and write
Note that, since both and are element-wise transformations, so is .
First, by using property (ii) of Prop. 4.6 (invariance of to element-wise transformation), we obtain
with such that is an isotropic Gaussian distribution.
Suppose for a contradiction that .
According to property (i) of Prop. 4.6, this entails that the matrix
Since by assumption, by property (i) of Prop. 4.6, the inner term on the RHS of (29),
is diagonal. Moreover, since is an element-wise transformation, and are also diagonal. Taken together, this implies that
is diagonal (i.e., (29) is of the form for some diagonal matrix ).
Without loss of generality, we assume the first component of is non-Gaussian and satisfies the assumptions stated relative to (axis not invariant nor sent to another canonical axis).
Now, since both the Gaussian CDF and the CDF are smooth (the latter by the assumption that of is a smooth density), is a smooth function, and thus has continuous partial derivatives.
By continuity of the partial derivative, the first diagonal element of must be strictly monotonic in a neighborhood of some (otherwise would be an affine transformation, which would contradict non-Gaussianity of ).
On the other hand, our assumptions relative to entail that there are at least two non-vanishing coefficients in the first row of (i.e., first column of ).In short, if this were not the case, this column would have a single non-vanishing coefficient, which would need to be one due to the unit norm of the rows of this orthogonal matrix. Such structure of the matrix would entail that the associated canonical basis vector is transformed by into a canonical basis vector which contradicts the assumptions. Let us call such pair of coordinates, i.e., and .
Now consider the off-diagonal term of (29), which we assumed (for a contradiction) must be zero almost surely w.r.t. . Since the term in (30) is diagonal, this off-diagonal term is given by:
where for the first equality we have used the fact that is a conformal map with conformal factor (by assumption), and where the second equality must hold almost surely w.r.t. .
Since is invertible, it has non vanishing Jacobian determinant. Hence, the conformal factor must be a strictly positive function, so
Thus, for almost all , we must have:
Now consider the first term in the sum.
Recall that , and that is strictly monotonic on a neighborhood of .
As a consequence, is also strictly monotonic with respect to on a neighborhood of (where the other variables are left constant), while the other terms in the sum in (31) are left constant because is an element-wise transformation.
This leads to a contradiction as (31) (which should be satisfied for all ) cannot stay constantly zero as varies within the neighbourhood of .
Hence our assumption that cannot hold.
We conclude that . ∎
Appendix D Worked out example
Consider the following example of a nonlinear ICA model which represents a change of basis from polar to Cartesian coordinates:
First, we consider the Jacobian of the true mixing which is given by:
and its determinant and column norms are given by
In other words, the columns of are orthogonal for all , so that for the true solution.
First, we write the joint density of using the change of variable formula:
Next, we compute the marginal density . Note that the observations live on the disk of radius , , so whenever .
Applying the change of variable with , and using the integral , as well as the fact that is an odd function, we obtain
Next, we compute the conditional density :
Finally, we compute the off-diagonal term in the general form of the inverse Jacobian for Damois-style solutions in (27):
Using the derivative and repeatedly applying the chain rule, we obtain:
Again, recall that this only holds inside the disk of radius , otherwise (as the CDF will be zero or one, irrespective of ).
The for the Darmois solution thus takes the form:
where the strict inequality in the last step follows from the fact that the fraction inside the logarithm, and hence the entire integrand, is strictly positive within the disk of integration.
We have thus shown that for the example of an orthogonal coordinate transformation from polar to Cartesian coordinates, which is not a conformal map, the os the true solution is zero and that of the Darmois construction is strictly greater than zero, hence the two can be distinguished based on the value of the contrast.
Appendix E Experiments
The code for our experiments (enclosed in the supplemental material) is in Python; we use Jax , Distrax and Haiku to implement our models; the Jacobian and computation and optimisation are performed with the automatic differentiation tools provided in Jax.
E.2 How to implement the Darmois construction
In the following, we describe how the Darmois construction can be implemented based on normalising flow models . The key idea is that the components of the Darmois construction (4) are conditional (cumulative) density functions corresponding to a given factorisation of the likelihood. A flow model with triangular Jacobian can be used to maximise the likelihood of the observations under a change of variable respecting said factorisation, and learning to map the observed variables onto a given (factorised) base distribution. After training, and provided that the model is expressive enough, the CDF of each component of the reconstructed sources should match that of the base distribution. By further transforming each reconstructed variable through said CDF, we achieve a global mapping of the observations onto a Uniform distribution on the -dimensional hypercube, with a triangular Jacobian, matching the transformation operated by the Darmois construction (see also see , section 2.2). Note that, for the purpose of computing the of the Darmois construction, this final step can be omitted due to Prop. 4.6, (ii), stating that the contrast is blind to element-wise reparametrisations of the sources.
We remark that, while the possibility of using normalising flows to “learn” the Darmois construction is mentioned in , where a similar construction is mentioned in a theoretical argument to prove “universal approximation capacity for densities” for normalising flow models with triangular Jacobian, it has to the best of our knowledge not been tested empirically, since autoregressive modules with triangular Jacobian are typically used in combination with permutation, shuffling or linear layers which overall lead to architectures with a non-triangular Jacobian.
To obtain an expressive normalizing flow with triagular Jacobian, we modify the residual flow model .We describe how to implement a function with upper triangular Jacobian, but the reasoning can be extended to implement functions whose Jacobian is lower triangular. A residual flow is a residual network which is made invertible through spectral normalization. Each layer is given by
Here, is the number of hidden units dedicated to transforming with the constraint . We perform an even split such that the and differ by at most 1 for . The weight matrices are restricted to be block triangular during optimization by setting the respective matrix elements to zero after each iteration of the optimizer. The model can simply be made and kept invertible using the same spectral normalization as is used for dense residual flows . We train our model to map onto a standard Normal base distribution.
E.3 Generating random MLP mixing functions
In order to generate random MLP mixing functions, we adopt the same initalisation as in : we initialise the square weight matrices to be orthogonal,Note that orthogonality of the weight matrices in a MLP does not guarantee satisfying Principle 4.1, due to the element-wise nonlinearities between the layers, which overall lead to a Jacobian whose columns are in general not orthogonal. and use the leaky_tanh invertible nonlinearity.
The modified maximum likelihood objective described in § 5.2 can be written as follows:while the objective in § 5.2 involves an expectation over , we consider the loss for a single point here, .
where represents the -th column of the inverse of the Jacobian of computed at .
We use the same model as the one described in § E.2, but without the constraint that the Jacobian should be triangular, and train with a Logistic base distribution.
Note that the computational efficiency of optimising objective (38) is cubic in the input size , due to a number of operations (matrix inversion, Jacobian and determinant computation via automatic differentiation, etc.) which are . However, similarly to what already observed in , we found that for data of moderate dimensionality computing and optimising objective (38) with automatic differentiation is feasible. For example, training a residual flow with 64 layers for iterations takes roughly 5.3 hours for , 5.7 hours for , and 6.3 hours for on the same hardware (see section E.5). An interesting direction for future work would be to find computationally efficient ways of optimising (38).
When computing the of the Darmois solutions of randomly generated functions, we restricted ourselves to Möbius transformations, i.e. conformal maps. However, there are also nonconformal maps satisfying , e.g. the transformation of Cartesian to Polar coordinates, see Appendix D. To test whether the of the Darmois solutions is actually bigger than , we gener
E.5 Evaluation
To evaluate the performance of our method, we compute the mean correlation coefficient (MCC) between the original sources and the corresponding latents, see for example . We first compute the matrix of correlation coefficients between all pairs of ground truth and reconstructed sources. Then, we solve a linear sum assignment problem (e.g. using the Hungarian algorithm) to match each reconstructed source to the ground truth one which has the highest correlation with it. The MCC matrix contains the Spearman rank-order correlations between the ground truth and reconstructed sources, a measure which is blind to nonlinear invertible reparametrisations of the sources.
While the MCC metric evaluates BSS by comparing ground truth and reconstructed sources, we propose an additional evaluation directly based on comparing the (Jacobians of the) true mixing and the learned unmixing. We take inspiration from an evaluation metric used in the context of linear ICA, the Amari distance : Given a learned unmixing and the true mixing , and defining the matrix , the Amari distance is defined as
and is greater than or equal to zero, canceling if and only if is a scale and permutation matrix, that is when the learned unmixing is matching the unresolvable ambiguities of linear ICA.
We extend this idea to the nonlinear setting: Given a true mixing and a learned unmixing , we define our nonlinear Amari distance as
Then, according to the definition of Amari distance (39), if the smooth function is a permutation composed with a scalar function, thus precisely matching the BSS equivalence class defined in Defn. 2.2, this would result in its Jacobian (that is, the product of the Jacobians ) equalling the product of a diagonal matrix and a permutation matrix at every point : the quantity would therefore be equal to zero.
This metric can be of independent interest and potentially useful in contexts where the reconstructed sources might be a noisy version of the true ones, but the true unmixing is nevertheless identifiable. Our implementation is based on the one for the (linear) Amari distance provided in the code for .
When computing the of the Darmois solutions of randomly generated functions, we restricted ourselves to Möbius transformations which are conformal maps. However, there are also nonconformal maps satisfying , e.g., the transformation from polar to Cartesian coordinates with , see Appendix D. To test whether the of the Darmois solutions is actually bigger than , we generate random radial transformations by imposing a random scale and shift before applying the radial transformation, compute the Darmois solution as we have done in § 5.1, and calculate its on the test set. We did 50 runs and the results are shown in Fig. 12.
Similar to Fig. 4 (a) we can clearly see that all values of the final models are larger than 0, with the smallest value being . This confirms the result we have already shown theoretically.
We show additional plots for the quantitative experiments involving training with the objective described in (38), see Fig. 8, Fig. 9 and Fig. 10.
For (that is, ground truth mixing linear), there appears to be an almost perfect recovery of the ground truth sources (resp. unmixing function) for , as can be seen by the high (resp. low) values of the MCC (resp. nonlinear Amari distance) evaluations ; this is in stark contrast with the distribution of the MCC (resp. nonlinear Amari distance) values for models trained with , which are typically much higher (resp. lower), indicating that the learned solutions do not achieve blind source separation (see , Fig. 8 (g), (h); , Fig. 9 (g), (h)). All models achieve a comparably good fit, reflected in the KL-divergence values (, Fig. 8 (e); , Fig. 9 (e)).
The trend is confirmed when the true mixing is nonlinear (), with slightly lower (resp. higher) values achieved with regularisation for the MCC (resp. nonlinear Amari) metrics; this possibly due to the increased difficulty of fitting observations generated by a nonlinear mixing, as can be seen from the higher values of the KL-divergence (, Fig. 8 (a); , Fig. 9 (a); , Fig. 10 (a));The distribution of the KL values contains outliers, and seemingly more strongly for lower values of . still, the beneficial effect of with respect to models trained with is clear, and is apparently stronger for and with higher data dimensionality (, Fig. 8 (c), (d); , Fig. 9 (c), (d); , Fig. 10 (c), (d)).
We additionally plot the values for the all trained models, for all values of . It can be seen that solutions found by unregularised maximum likelihood estimation typically learn functions with relatively high values of , while as expected the regularised version achieves low values (, Fig. 8 (b), (f); , Fig. 9 (b), (f); , Fig. 10 (b)).
Finally, in figure 11, we report the same plot as in 4, top row, but with a perceptually uniform colormap.
We compared the performance of our proposed regularised maximum likelihood procedure to a state of the art method for linear ICA, FastICA , in the implementation from the Scikit-learn package , over repetitions. Our experiments show that our regularised method (, and particularly ; provides the unregularised nonlinear baseline) is superior in learning the true unmixing and reconstructing the sources. This indicates that the linearity assumption of FastICA does not allow enough flexibility to solve blind source separation in our setting, whereas our criterion does (see Fig. 13, Fig. 14 and Fig. 15).the experimental setting and the plots for the normalising flow models correspond to those already shown in the paper, but here we modified the -axis scale to facilitate the comparison of all methods While the spread in the distributions of MCC and Amari distance can be largely attributed to the brittleness of neural networks, the median values for the MCC (resp. nonlinear Amari distance) are consistently higher (resp. lower) for our regularised method than for FastICA. In contrast, the performance of FastICA is consistently better than the unregularised baseline.
All models were trained on compute instances with 16 Intel Xeon E5-2698 CPUs and a Nvidia Geforce GTX980 GPU. The cluster we used has 204 thereof. Training the models took between 4 and 16 hours depending mainly on the dimensionality and number of samples in the dataset, and on the number of iterations used for training. Overall, we trained around 2000 models, amounting to roughly 18000 GPU hours.
Appendix F Additional background on conformal maps and Möbius transformations
A similarity of a Euclidean space is a bijection from the space onto itself that multiplies all distances by the same positive real number , so that for any two points and we have
where is the Euclidean distance from to . The scalar is sometimes termed the ratio of similarity, the stretching factor and the similarity coefficient. When a similarity is called an isometry (rigid transformation). Two sets are called similar if one is the image of the other under a similarity.
Note that such a similarity has Jacobian for any .
A characterisation of conformal maps directly related to orthogonal coordinate systems is the following.
While it can be shown that linear conformal maps are similarities, an interesting class of nonlinear conformal maps are the unit radius sphere inversion (restriction to unit radius is only to avoid unnecessary notational complexity):
We can notice that such transformation leaves the hypersphere of center and radius 1 invariant, while the points outside of the unit ball are mapped to the interior of the unit ball, and vice-versa.
Interestingly, conformal maps in Euclidean spaces of dimension superior or equal to 3 can be restricted to two kinds according to the following result from Liouville.
The class of function described in Thm. F.2 corresponds exactly to the Möbius transformations described in (32). These transformation can as well be defined in dimension , with the specificity that they are only a subset of the class conformal maps in this dimension.
We characterize the properties of the unit sphere centered at zero, that we denote
Now let us derive the Jacobian of . A straightforward computation leads to
where denote the identity matrix.
By noticing that is rank one symmetric with eigenvalue 1 associated with unit norm eigenvector , we can diagonalize this matrix in any (space dependent) orthogonal basis that has as the first basis vector.
Let us thus pick the unit vectors associated to the hyperspherical coordinates (which satisfy this condition by definition), and consider the orthogonal matrix gathering these basis vectors as its columns (it is parameterized by the unit vector , as this basis is radially invariant. Then we can write
with a diagonal matrix with diagonal elements . This leads to
with a diagonal matrix with diagonal elements . The Jacobian thus takes the form predicted by the above proposition for conformal maps
with scale factor and a space dependent orthogonal matrix, which has the additional property to be radially invariant for the specific case of sphere inversions.