Function Classes for Identifiable Nonlinear Independent Component Analysis
Simon Buchholz, Michel Besserve, Bernhard Schölkopf
Introduction
Unsupervised representation learning methods can fit Latent Variables Models (LVM) to complex real world data. While those latent representations allow to create realistic novel samples or represent the data in a compact way , they are a priori not related to the underlying ground truth generative factors of the data. It is highly desirable to recover the true underlying source distribution because those are expected to help with various downstream tasks, e.g., out of distribution generalization .
Several additional assumptions were suggested to make the ICA problem identifiable. Broadly, there are two directions. First, some works imposed additional or different restrictions on the distribution of the sources. One line of research adds temporal structure by considering time series data . More recently, Hyvärinen et al. proposed to introduce an observed auxiliary variable , e.g., a class label, such that the source distribution has independent components conditional on the auxiliary variable. They show that under suitable assumptions on the distribution of and arbitrary nonlinear mixing functions can be identified. Several recent works extended this approach .
Another possibility is to restrict the class of admissible functions by considering more flexible classes than just linear functions, but not allowing arbitrary non-linear functions. The general aim of this approach is to find sufficiently “small” function classes such that ICA is identifiable within this class while making them as large as possible to allow flexible representation of complex data and being applicable to real world problems. So far results in this direction are rather limited. It was shown that the post-nonlinear model is identifiable . Moreover, it has been shown that ICA with conformal maps in dimension is almost identifiable . More recently, identifiability of volume preserving transformations was investigated in the auxiliary variable case in (combining the two possible restrictions) and identifiability based on sparsity of the mixing function was studied in .
In this work, we extend the previous works by proving new identifiability results for unconditional ICA. Our main focus is conformal maps (i.e., maps that locally preserve angles) and Orthogonal Coordinate Transforms (OCT) (i.e., maps satisfying that is a diagonal matrix where denotes the derivative of ). OCTs, that we will also call orthogonal maps for simplicity, were recently introduced in the context of representation learning in , motivated by the independence of mechanisms assumption from the causality literature. The main focus of this work is to prove new identifiability and partial identifiability results for this class of functions. Our main contributions are the following.
We prove that ICA with conformal maps is identifiable in and extend the earlier results in dimension 2 (Theorems 2-3).
We define a new notion of local identifiability (Definition 5) and prove that ICA with orthogonal maps is locally identifiable (Theorem 4).
On the contrary we show that ICA with volume preserving maps is not identifiable not even in the local sense (Theorem 6).
We introduce new tools to the ICA field: our results are based on connections to rigidity theory, restricting the global structure of functions based on local restrictions. Moreover, in contrast to most earlier results that argue locally using results from linear algebra we exploit the global structure of partial differential equations related to the identifiability problem.
The remainder of this paper is structured as follows. In Section 2 we introduce the general setting of ICA and identifiability. Then we discuss our results for different classes of nonlinear maps. We consider conformal (Section 3), orthogonal (Section 4) and volume preserving (Section 5) maps. An overview of our results can be found in Table 1. Finally, in Section 6, we discuss the relation of identifiability of ICA to the rigidity of the considered function class.
Setting
We illustrate Definition 1 through the well known example of linear maps
This result is optimal as the ordering and scale of the is unidentifiable and the restriction to at most one Gaussian component is required to avoid linear MPTs of multivariate Gaussians. We provide a proof of this result in Appendix D as this serves as a preparation for the more involved Theorem 2 below.
Results for conformal maps
Our first main result is an extension of Theorem 1 to conformal maps. A conformal map is a map that locally preserves angles, i.e. locally it looks like a scaled rotation. It can be shown that this is equivalent to the following definition.
All our results also hold for the more general class of conformal maps where is a Riemannian manifold. The complete definition can be found in Appendix B. For convenience we define signed permutation matrices by
While this condition might appear a bit technical it actually only rules out pathological cases like the cantor measure or densities which are nowhere differentiable and probably it could be relaxed further. In particular contains all probability measures with piecewise smooth densities. Then the following identifiability for conformal maps in dimension holds.
This means that we can identify conformal maps up to three symmetries, namely constant shifts of the distributions, rescaling of all coordinates by the same constant factor, and permutations of the coordinates. The proof is in Appendix E. The main ingredient in the proof is that conformal maps in dimension are very rigid and can be characterized explicitly, as we will discuss in Section 6.
We remark that it might be more natural to not fix the scale of the sources and allow arbitrary coordinate-wise rescalings. The result can be easily extended to accommodate this. We define
The additional restriction on the observational distribution is clearly necessary to exclude the non-identifiability of Gaussian distributions.
Results for orthogonal maps
Recently, in , the more general class of OCTs was considered in the context of ICA. They referred to orthogonal coordinates as IMA maps, referencing to independent mechanisms. This nomenclature was motivated by the causality literature and we refer to their paper for an extensive motivation and further results. As we focus on theoretical results for this function class we stick to the more common term of OCTs. Orthogonal coordinate transformations are defined as the set of functions whose derivative have orthogonal columns, i.e., the vectors and are orthogonal for .
We first note that just as for linear maps and conformal maps (see Thm. 1) Gaussian distributions can hamper identification. This is because arbitrary measures with a factorized density can be pushed forward into multivariate Gaussians using a suitable coordinate-wise transformation.
An illustration of this definition can be found in Figure 3. Based on invariant deformations we can now define a local identifiability property of ICA in a given function class.
The proof, in Appendix F, introduces new tools to the field of ICA. The main idea is to consider the vector field that generates the deformation and then rewrite the assumption as systems of partial differential equations for . The proof is then completed by showing that the only solution of this system vanishes. Let us state one simple consequence of this theorem.
Let us reiterate what this corollary shows: we cannot smoothly and locally transform the function such that (1) the observational distribution remains invariant, i.e., equal to , and (2) the deformed functions remain OCTs.
At a high level this result suggests that OCTs can be identified if we know close to the boundary of the support of , e.g., by having, in addition to unlabelled data , labelled data for those where one coordinate is extremal. Note that we actually do not show this result as there might be further solutions which are not connected by smooth transformations. We expect that those results can be generalized substantially. In particular, we conjecture that for “most” functions the boundary condition can be removed thus giving a stronger local identifiability result up to the boundary of the support of . As a partial result in this direction we prove the following theorem.
In Appendix F we show that this theorem follows from a slightly stronger result stated as Theorem 11 which has a similar proof as Theorem 4. We do expect that the conclusion of the Theorem actually holds for all not just almost all, but we are unable to show this.
For completeness, we provide a general construction that is close to our proof of Theorem 4 in Appendix C in the supplement. A very clear construction for this result was given in .
An illustration of this construction is shown in Figure 3 (see App. C for details). Clearly, by concatenation this allows us to create a vast family of spurious solutions. Note that those solutions are excluded when restricting to OCTs which is a corollary of Theorem 4.
Results for volume preserving maps
The proof and an illustration are in Appendix G. It is based on the flows generated by suitable explicit vector fields. As those flows can be concatenated we obtain a large family of spurious solutions. We think that the approach used here is a powerful technique to construct counter-examples to identifiability in ICA. Note that while is not identifiable it can be possible to identify certain values of when we the distribution of is known, e.g., volume preservation implies that the source value with the largest density is mapped to the point with the largest density in the observational distribution. As local identifiability is weaker than identifiability we get the following (informal) corollary.
Relation to rigidity theory
Our results rely on rigidity properties of certain function classes. Rigidity refers to the property that a local constraint on the derivative of a function implies global restrictions on the shape of the function. For conformal maps the following rigidity result holds. Let us clarify this through a well known example (a special case of Theorem 8 below).
One important result in statistical mechanics is that results like Theorem 7 come with estimates in a neighbourhood of the function class, in the sense that if is close to a rotation for all , then it will be close to a constant rotation . Results of this type could be important for three reasons. First, in many real world applications it might be more realistic to assume that is close to a certain function class but not necessary contained in .
Secondly, it could rigorously justify the use of surrogate losses for the differential constraints (e.g., consider the loss ) We denote the Euclidean norm by . and, thirdly, results of this type would be needed for finite sample analysis.
Originally this result was shown by Liouville , a modern treatment is . In particular, this shows that conformal maps are (up to translations) rotations or rotations followed by an inversion.
To illustrate the strength of Theorem 8 we compare it to the setting of volume preserving maps which satisfy no similar rigidity property. Intuitively the different rigidity properties are already apparent from the connection to solids, which can merely be rotated and shifted, and fluids which can also be stirred leading to chaotic deformations. Rigorously the different behaviours can be clarified by the observation that conformal maps have a finite number of parameters and thus a finite number of constraints (e.g., of the form ) allows to identify them. In contrast volume preserving maps cannot be identified from finitely many constraints as the following proposition shows.
Let us emphasize that the different rigidity properties are not at all surprising when arguing based on degrees of freedom or numbers of constraints. While volume preserving maps enforce only a single scalar constraint on the Jacobian the condition for conformal maps gives constraints on the Jacobian.
Let us finally comment on OCTs where the picture is not as well understood. As discussed in Section 4 it is known that OCTs constitute a rich, non-parametric class of functions and therefore OCTs are much more flexible than conformal maps. We illustrated this with example OCT constructions in Section 4 leveraging 2D conformal maps. Nevertheless, it is not known if and what rigidity properties can be derived for OCTs. However, our results suggest that the additional measure preservation condition in the context of ICA gives enough rigidity to (almost) give identifiability of ICA. In this sense OCTs might be a good function class for ICA as it is rich enough to allow complex representations of data while at the same time being sufficiently rigid to still provide a notion of identifiability whose strength remains to be determined.
Discussion
ICA is long known to be identifiable for linear maps, baring pathological cases, and highly non-identifiable for general nonlinear ones. Surprisingly, similar results for function classes of intermediate complexity remain scarce. In this work we address this question with several identifiability results for different function classes. Our first main result is that ICA is identifiable in the class of conformal maps (up to classical ambiguities). This considerably extends previous claims, limited to a specific 2D setting , and ruling out several families of spurious solutions . On the negative side we show that the ICA problem for volume preserving maps admits a large class of spurious solutions. Finally, we show that OCTs satisfy certain weaker notions of local identifiability.
In our proofs, we draw connections to methods and techniques that, to the best of our knowledge, have not been used in the context of ICA before. We relate the identifiability problem in ICA to the rigidity of the considered function class and use tools from the theory of partial differential equations. These techniques have been applied very successfully to the analysis of elastic solids and we believe that there are many applications of these methods in the field of ICA.
While the main focus of current research after the seminal work of Hyvärinen et al. is on the auxiliary variable case, there are three reasons to consider unconditional ICA. Firstly, it is a fundamental research question that is, as illustrated by our results, deeply rooted in functional analysis. Secondly there is high application potential for completely unsupervised learning without any auxiliary variables, as the corresponding datasets do not require labelling or specific experimental settings. Thirdly, it is very likely that the techniques can be generalised to the auxiliary variable case.
Finally, a central question from a machine learning perspective is the ability to design learning algorithms that can train LVMs with identifiable function class constraints. Interestingly, Gresele et al. showed that OCT maps can be learnt using a closed from regularized likelihood loss, thereby providing, supported by our result, a full-fledged identifiable nonlinear ICA framework.
References
Appendix
In the appendix we provide the proofs of our results and we discuss the relevant background. It is structured as follows. We first introduce some mathematical background in Appendix A and extend the function class definitions to Riemannian manifolds in Appendix B. We discuss a general construction of spurious solutions in Appendix C. Then we provide the proofs of our main results in Sections E to G.
For the convenience of the reader we collect some mathematical definitions and notations. All definitions can be found in standard textbooks.
For a measure on a (measurable) space and a measurable map, the pushforward measure is defined by for any measurable set . Here denotes the preimage of under . Sometimes the pushforward measure is denoted by . Note that no further restrictions on are necessary, in particular does not need to be invertible. One important property that we will use frequently is the relation .
Note that the potentially more familiar version for random variables reads as follows. Let and be random variables satisfying , then their densities are related by
Let us remark concerning the notation that it is convenient to put the argument below, i.e., we write . Moreover, when applying differential operators they will by default only act on the spatial variable , i.e., denotes the derivative of the function for a fixed . Note that the solutions of differential equations can exhibit blow-up so might not be defined for all times . However, when we assume that is bounded the flow exists globally.
We at some places use the notation . We also use the notation. Recall that as means that there are constants and such that for . Recall that we introduced the notation in the main part and we denoted by the uniform measure on . We write iff as a shorthand for ’if and only if’.
Appendix B Function class definitions for general manifolds
where denotes the pullback metric. This means that for
Moreover, we observe that the adjoint satisfies by definition
i.e., has orthogonal columns with equal norm. Note that the concatenation of conformal maps is conformal and the inverse of conformal diffeomorphisms is again conformal.
Again, this condition can be equivalently written as
An important remark is that orthogonal coordinates do not exist for all manifolds as there are obstructions. Manifolds with this property are called locally diagonalizable. This is closely related to the representation capability of the function class.
Appendix C Measure preserving transformations and spurious solutions
and similarly the subset of left-composable functions
Appendix D Proof of Theorem 1
We here, for completeness, give a proof of Theorem 1. While this result is well known we think that it makes sense to include a condensed proof because it contains many of the key steps of the more involved proof for conformal maps in the next section and it is not as well known as the proof based on Darmois-Skitovich Theorem which does not generalise to nonlinear functions.
The main idea of the proof is to use the observation that for and all such that
i.e., mixed second derivatives of the log density vanish. We plug (28) into this equation and get
We now denote . Then we get (using that is invertible) that for all such that
As is not a Gaussian for and thus not constant we conclude
for and all . Plugging this into (32) we obtain
Note that if is a Gaussian density then so we conclude that in any case for , i.e., (33) holds as well for .
Let us clarify what happens when there is more than one Gaussian component. In this case there might be multiple constant non-zero terms in (32) whose contributions can cancel and we cannot conclude that (33) holds for all . This recovers the well known non-uniqueness for Gaussian variables.
Appendix E Proofs for the results on conformal maps
In this section we give the proofs for Section 3. First, we consider and then the special case .
The proof of Theorem 2 uses similar ideas as the proof for Theorem 1, however, the calculations are more involved. The key ingredient is the classification of all conformal maps in Theorem 8. From there we see that we already dealt with the linear case in Theorem 1, so it is sufficient to focus on the case of nonlinear transformations. Recall that the Moebius transformations introduced in (11) are given by
Let us quickly show how it implies Theorem 2 before we prove this lemma.
To prove Lemma 2 we need one technical result that shows that local properties of the density (i.e., properties that hold for for some non-empty open sets ) in fact hold for all . This will be based on the nonlinearity of the map combined with the factorisation of the densities. An illustration can be found below in Figure 4.
In Step 1 and 2 we eliminate the trivial symmetries of the measure and show that the mild local regularity assumption on the measures imply global regularity.
Then, in Steps 3 and 4 we derive in (47) a condition similar to (32) but more involved.
To exploit this condition we look in Steps 5 and 6 at certain limiting regimes where the terms become much simpler and almost reduce to the condition of the linear case. This allows us to conclude that is a permutation matrix in Step 7.
Using that is a permutation matrix in (47) we get in Step 8 a much simpler relation that almost factorizes. This allows us to derive a simple differential equation in Step 9 which restricts the potential densities to a simple parametric form. This allows us to conclude.
Let us emphasize here, that we already finished the proof of Lemma 2 for probability distributions with bounded support. This is more difficult in dimension 2 because there are two-dimensional conformal functions mapping rectangles to rectangles. We will consider this in the proof of Theorem 3 below.
A major step in the proof is to show that is a permutation matrix under the assumptions of the lemma. We define the index set as the set of all indices such that the -th row of the matrix has only one non-zero entry. Our goal is to show that . We have shown so far that is positive and twice differentiable if .
The proof in the linear case relied on the relation (32). We now derive a similar relation for non-linear Moebius transformations.
We apply the same reasoning as in the proof of Theorem 1 to derive partial differential equations for the density . The condition , or equivalently and the density relation (13) imply
Evaluating the derivatives using (39) and (40) we get
Finally we express the variable through . Note that then and . Plugging this in the last equation we get
where we used as is orthogonal and we used Einstein summation convention to sum over indices that appear twice (we kept the sum over for better readability). Note that this expression is not homogeneous in .
By varying this almost implies that for some constant whenever and twice differentiable. However, for this we need to show that the terms hidden in are really negligible, i.e. and are bounded as so that the remainder becomes which we will establish.
Note that if such a relation could be derived we could conclude, similarly to the linear case, that is a permutation matrix.
Recall that is the set of indices such that the -th row of has only one non-zero entry. Let and pick such that . Fix all coordinates except so that is positive and twice differentiable at . Then we can express (49) as
where the remainder term contains the remaining terms. The expression of course depends on the other coordinates for but since they are considered fixed here we can view as a function of alone.
Equation (49) then implies that there is sufficiently large such that for the remainder term satisfies for some constant .
Here the last constant term bounds the contribution. Suppose . Then we can bound
We find that (for this is clear, and otherwise we can absorb the absolute value part). We conclude by integration that for all (note that the bound is trivially true if )
Similarly we can bound for
implying for such that . We obtain
Together the last two steps imply that for some and . Going back to (50) we conclude that there is such that for . We conclude that for and all
If we are done. So there is and using (56) we conclude by varying that
for some constant (depending on , , and ). Note that if we assumed that at most one is a Gaussian density we could conclude as in the linear case. However, this assumption is not necessary, as we will now show.
Dividing (59) by and subtracting it from (58) we conclude that
We now assume that . Let , , and be pairwise different. Using the last display with , and , and subtracting the resulting equations we obtain
Varying and independently and since is arbitrary we conclude that there is a constant such that
The solutions of the ODE are given by
where is any constant. We conclude that there are constants such that
By applying the main argument to we infer that has to have again the same structure as in (66) so we conclude that and . Alternatively, one directly sees that only factorizes as if those conditions hold. It is easy to see that those densities satisfy the assumptions. However, is never integrable so there are no probability distributions satisfying the relations This ends the proof for .
For we cannot simplify (62) by considering indices . Instead, we directly exploit (62) to obtain a similar conclusion. Similarly to the argument in Step 6 it can be shown that and are bounded for away from 0. Then we consider in (62) and divide by . We get (using )
By varying we conclude just as for that is constant. We conclude as before. ∎
Let us now prove the geometric result from Lemma 3 above.
The main idea of the proof is that a box contained in after inversion is distorted so that its convex hull (contained in ) is strictly bigger than the box image so that inverting backwards gives us a bigger box in except for some special cases. An illustration of this argument is shown in Figure 4. The formal argument below is slightly technical. An illustration of the actual argument can be found in Figure 5.
W.l.o.g. we now suppose that there is a box with . We write , . We consider the point
We bound (using )
Let be maximal such that . Then the reasoning above shows that . The same reasoning for the other coordinates implies that . By applying the same reasoning to sequences of boxes in approaching the origin we conclude that is the union of quadrants (and the same holds for ).
It remains to prove the last remark. As quadrants are invariant under we have and conclude , or equivalently . It is sufficient to show that . For simplicity we assume , the generalisation to other quadrants is immediate. Since we assume that the -th row of is not equal to a signed standard basis vector it has at least two non-zero entries. Thus we can find such that and all entries of are non-zero. Since is orthogonal to the span of there is a vector such that and . By adding a suitable vector we can ensure that , all entries of are non-zero and for . The second condition can be satisfied by picking the entries of one after another. The conditions for and imply that since we assumed . But then and since is strictly contained in a quadrant (all entries are non-zero) we conclude that and thus . This implies .
E.2 Proof of Corollary 1
Here we prove the simple extension of Theorem 2 to rescaled conformal maps.
E.3 Proof of Theorem 3
where denotes the complex derivative. The proof of Theorem 3 is based on the fact that conformal maps between rectangles can be characterized rather explicitly. In particular, we use the following result.
A proof of this result can be found in any textbook on complex analysis, e.g., . With this result we can prove Theorem 3.
where and we used that all angles in a rectangle are . The points are the preimages of the corners of . Then the derivative of can be written as
Thus so we conclude that
Let us finally remark that the identifiability of conformal maps for distributions with full support in follows just as in because every bijective conformal map of the Riemann sphere to itself is a Moebius transformation so we can apply Lemma 3. We expect that the result can be extended to more general densities using the same strategy and a more careful analysis of the density close to the boundary.
Appendix F Proofs for the results on OCTs
In this section we collect the missing proofs for Section 4.
First we prove Proposition 1. We refer to Appendix C for a general review and characterisation of spurious solutions.
The polar coordinates defined above satisfy:
The determinant of the Jacobian is given by
For the first part we only need to show that has orthogonal columns which can formally be done by induction noting that
The determinant and the image can also be derived from this recursion. ∎
Clearly . They are strictly increasing functions with positive derivative, i.e., diffeomorphisms, so we can define their inverses on an open interval which are also differentiable functions. Note that
where is the volume of the unit ball in dimension and these constants ensure that is a probability density. Define as the cdf for the probability density , i.e.,
Since iff we conclude that restricted to is a continuous and strictly increasing from 0 to 1 and has a positive derivative. Hence, we can define by and is differentiable with
We define the domains Now we consider the map given by
Note that is a coordinate-wise transformation and the determinant of its Jacobian is given by
Note that by definition of we have and acts coordinatewise where the action on the first coordinate is so that
We conclude, using (89) and the last display, that
F.2 Proof of Theorems 4
We now consider smooth deformations of a data generating mechanism . For this it is helpful to phrase these as flows generated by vector fields. For a brief review of these notions we refer to Appendix A and for an extensive introduction we refer to any textbook on differential geometry. We now give a complete proof of Theorem 4.
We now evaluate in terms of and . To evaluate the time derivative it is convenient to write instead of . We get, using ,
Combining this with the last display we get (dropping the positional argument for conciseness)
Now we use that and map to diagonal matrices for all , in particular and are diagonal matrices. We conclude that for the equation
holds. Thus, we obtain a system of first order Partial Differential Equations (PDE) for . We now write . We also fix an in the following. Then we can rewrite (109) concisely as
We divide equation (110) by apply and sum over to obtain
This implies that satisfies the wave equation
We now consider as above and get using the last display
Plugging the relations in (123) we obtain
Thus we established that the relation (107) also holds in the manifold case. The rest of the proof is the same. ∎
Apply Theorem 4 with constant. The assumption that is analytic in can be dropped as explained in the proof of Theorem 4. ∎
where we assume , , and that there is such that
Under the assumptions above there is a unique weak solution of the system (126) with boundary values as in (127) and (128).
For our purposes it is not necessary to define weak solution let us just emphasize that any classical solution is a weak solution so this implies uniqueness of classical solutions.
The key obstacle to improve upon this result and to remove the compact support condition on is that the resulting PDE in equation (112) is well posed for the Cauchy initial value problem but it is not well posed for the Dirichlet problem or for mixed Dirichlet and Neumann boundary data. In particular, solutions are, in general, not unique. Furthermore, there are no general uniqueness results for first order systems as in (109). Note that the existence of a non-trivial divergence free solution of (109) does not imply that a non-constant flow exists because this is not sufficient to define the flow for positive times.
We illustrate the influence of the boundary condition further below, when we prove Theorem 5.
F.3 Proof of Theorem 5
In this section we show that a family of simple mixing functions is locally identifiable for most parameter values even when the mixing is not known close to the boundary. Note that actually we can construct a set of parameter values for which this holds giving a slightly stronger result that we state now. Theorem 5 will be simple consequence of this result.
The initial part of the proof proceeds as in the proof of Theorem 4 and we keep using the same notation. In particular and is given by .
Note, that now is constant in and therefore is constant in . We now investigate the boundary conditions for equation 112. Note that
So preserves and we conclude that . Let us denote by
the boundary hyperplanes and write . As maps bijectively to itself we conclude that
We now focus on and use the shorthand . Then the differential equation (110) implies that
We conclude that the function solves the following mixed Dirichlet and Neumann type boundary problem
in particular is constant. So the equation (135) becomes a constant coefficient hyperbolic equation which can be solved explicitly.
We can now use Theorem 1 from (and a simple scaling argument) we conclude that the system (135) has a unique solution which is (actually this result is for on but the proof is still valid). To give an intuition, we note that separation of variable is possible in this setting and all solutions to the boundary value problem (135) and (136) (i.e., without the boundary condition (137) for ) can be expressed as a linear combination of the form
Using now that (by (137)) we conclude and therefore
where , or equivalently
The complete argument goes as follows. We take the time derivative of equation (108) (recall that as and get, denoting and ,
and . The same arguments as before imply on . By induction all time derivatives of vanish. This implies that for all and , i.e., its Taylor expansion at disappears and since we assumed to be analytic in so is and we conclude that and therefore . ∎
F.4 Proofs for the construction of spurious solutions
Finally, we show how flows can be used to construct families of solutions to the ICA problem. This section contains the technical results missing in the overview given in Appendix C.
The first construction was described in Lemma 1. Let us for completeness give a proof (we emphasize again that this result is essentially taken from ).
Note that it is sufficient to show that the maps are volume preserving for fixed so we ignore the time argument. It is easy to see that is bijective (the inverse is given where ). Then we only need to show that for all . We calculate (denoting )
We conclude (writing )
Then we obtain, using the matrix determinant lemma for rank 1 updates (
We have therefore shown , completing the proof. ∎
Then we get . So those vector fields are divergence free and we conclude that the space
is infinite dimensional. Every generates a flow defined by
Appendix G Proofs for the result on volume preserving maps
Next we show that this construction can be generalised to volume preserving transformations and we prove Theorem 6. Note that in the special case that the distribution of is the construction above already works. This is a special case because the condition already implies that is volume preserving as soon as is volume preserving as the density of is constant. So in this case the condition that is volume preserving and essentially agree which is not the case for general base measures.
The constructed flows are non-trivial, i.e., not constant because the probability density cannot be constant (as we assumed it to be ) and thus is not identically vanishing.
It is easy to see (e.g., through the example above) that the flows will, in general, mix the coordinates and thus this really shows that ICA is not identifiable for volume preserving maps.
While we construct a finite family of solutions they can be combined, e.g.,
Let us finally sketch a proof of Proposition 2.
Appendix H Experimental illustration of local identifiability
We work in dimension 2 and consider a standard normal base distribution . We consider polar coordinates for
Note that we shift the first coordinate and scale the angular coordinates to ensure that the map is injective (except for the light tail of the Gaussian). We also consider the setting where
To measure how well the model recovers the ground truth sources we consider
For training we use stochastic gradient descend with the ADAM-optimizer and train for 1000 steps with a batch size of 256 where we generate i.i.d. samples from the observational distribution in each step. We averaged our results over 10 runs. Total compute time was less than 24h on a workstation.