Information Theoretical Estimators Toolbox
Zoltan Szabo
Introduction
Since the pioneering work of Shannon 1948, entropy, mutual information, The Shannon mutual information is also known in the literature as the special case of total correlation or multi-information when two variables are considered. association, divergence measures and kernels on distributions have found a broad range of applications in many areas of machine learning (Beirlant et al. 1997; Wang et al. 2009; Villmann and Haase 2010; Basseville 2013; Póczos et al. 2012; Sriperumbudur et al. 2012). Entropies provide a natural notion to quantify the uncertainty of random variables, mutual information and association indices measure the dependence among its arguments, divergences and kernels offer efficient tools to define the ‘distance’ and the inner product of probability measures, respectively.
A central problem based on information theoretical objectives in signal processing is independent subspace analysis (ISA; Cardoso 1998), a cocktail party problem with independent groups. One of the most relevant and fundamental hypotheses of the ISA research is the ISA separation principle (Cardoso 1998): the ISA task can be solved by ICA (ISA with one-dimensional sources, Hyvärinen et al. 2001; Cichocki and Amari 2002; Choi et al. 2005) followed by clustering of the ICA elements. This principle (i) forms the basis of the state-of-the-art ISA algorithms, (ii) can be used to design algorithms that scale well and efficiently estimate the dimensions of the hidden sources, (iii) has been recently proved (Szabó et al. 2007) and (iv) can be extended to different linear-, controlled-, post nonlinear-, complex valued-, partially observed systems, as well as to systems with nonparametric source dynamics. For a recent review on the topic and ISA applications, see Szabó et al. 2012.
Although there exist many exciting applications of information theoretical measures, to the best of our knowledge, available packages in this domain focus on (i) discrete variables, or (ii) quite specialized applications/information theoretical estimation methods. Our goal is to fill this serious gap by coming up with a (i) highly modular, (ii) free and open source, (iii) multi-platform toolbox, the ITE (information theoretical estimators) package, which focuses on continuous variables and
is capable of estimating many different kind of entropy, mutual information, association, divergence measures, distribution kernels based on nonparametric methods. It is highly advantageous to apply nonparametric approaches: the ‘opposite’ plug-in type methods—estimating the underlying densities—scale poorly as the dimension is increasing.
offers a simple and unified framework to (i) easily construct new estimators from existing ones or from scratch, and (ii) transparently use the obtained estimators in information theoretical optimization problems,
with a prototype application in ISA and its extensions.
Library Overview
Below we provide a brief overview of the ITE package:
The ITE toolbox is capable of estimating numerous important information theoretical quantities including
Shannon-, Rényi-, Tsallis-, complex-, -, Sharma-Mittal entropy,
generalized variance, kernel canonical correlation analysis, kernel generalized variance, Hilbert-Schmidt independence criterion, Shannon-, -, Rényi-, Tsallis-, Cauchy-Schwartz quadratic-, Euclidean distance based quadratic-, complex-, mutual information; copula-based kernel dependency, multivariate version of Hoeffding’s , Schweizer-Wolff’s and , distance covariance and correlation, approximate correntropy independence measure,
Kullback-Leibler-, -, Rényi-, Tsallis-, Cauchy-Schwartz-, Euclidean distance based-, Jensen-Shannon-, Jensen-Rényi-, Jensen-Tsallis-, K-, L-, Pearson -, f-divergences; Hellinger-, Bhattacharyya-, energy-, (non-)symmetric Bregman-, J-distance; maximum mean discrepancy,
multivariate (conditional) extensions of Spearman’s , (centered) correntropy, correntropy induced metric, correntropy coefficient, centered correntropy induced metric, multivariate extension of Blomqvist’s , lower and upper tail dependence via conditional Spearman’s ,
expected-, Bhattacharyya-, probability product-, (exponentiated) Jensen-Shannon-, (exponentiated) Jensen-Tsallis-, exponentiated Jensen-Rényi kernel.
ITE offers solution methods for independent subspace analysis (ISA) and its extensions to different linear-, controlled-, post nonlinear-, complex valued-, partially observed systems, as well as to systems with nonparametric source dynamics; combinations are also possible. The solutions are based on the ISA separation principle and its generalizations (Szabó et al. 2012).
Beyond IPA, ITE provides quick tests to study the efficiency of the estimators. These tests cover (i) analytical value vs. estimation, (ii) positive semi-definiteness of Gram matrices defined by distribution kernels and (iii) image registration problems.
The core idea behind the design of ITE is modularity. The modularity is based on the following four pillars:
The estimation of many information theoretical quantities can be reduced to k-nearest neighbor-, minimum spanning tree computation, random projection, ensemble technique, copula estimation, kernel methods.
The ISA separation principle and its extensions make it possible to decompose the solutions of the IPA problem family to ICA, clustering, ISA, AR (autoregressive)-, ARX- (AR with exogenous input) and mAR (AR with missing values) identification, gaussianization and nonparametric regression subtasks.
Information theoretical identities can relate numerous entropy, mutual information, association, cross- and divergence measures, distribution kernels (Cover and Thomas 1991).
ISA can be formulated via information theoretical objectives (Szabó et al. 2007):
where the minimizations are w.r.t. the optimal clustering () of the ICA elements.
The ITE package offers dedicated solvers for the obtained subproblems detailed in ‘Modularity:1-2’. Thanks to this flexibility, extension of ITE can be done effortlessly: it is sufficient to add a new switch entry in the subtask solver.
We illustrate how easily one can estimate information theoretical quantities in ITE:
Next, we demonstrate how one can construct meta estimators in ITE. We consider the definitions of the initialization and the estimation of the J-distance. The KL-divergence, which is symmetrised in J-distance, is estimated based on the existing k-nearest neighbor technique.
Due to the unified syntax of the estimators, one can formulate and solve information theoretical optimization problems in ITE in a high-level view. Our example included in ITE is ISA (and its extensions) whose objective can be expressed by entropy and mutual information terms, see ‘Modularity:4’. The unified template structure in ITE makes it possible to use any of the estimators (base/meta) in these cost functions.
A further attractive aspect of ITE is that even in case of unknown subspace dimensions, it offers well-scaling approximation schemes based on spectral clustering methods. Such methods are (i) robust and (ii) scale excellently, a single general desktop computer can handle about a million observations—in our case estimated ICA elements—within several minutes (Yan et al. 2009).
Availability and Requirements
The ITE package is self-contained, it only needs a Matlab or an Octave environment See http://www.mathworks.com/products/matlab/ and http://www.gnu.org/software/octave/. with standard toolboxes. ITE is multi-platform, it has been extensively tested on Windows and Linux; since it is made of standard Matlab/Octave and C++ files, it is expected to work on alternative platforms as well. On Windows (Linux) we suggest using the Visual C++ (GCC) compiler. ITE is released under the free and open source GNU GPLv3 () license. The accompanying source code and the documentation of the toolbox has been enriched with numerous comments, examples, detailed instructions for extensions, and pointers where the interested user can find further mathematical details about the embodied techniques. The ITE package is available at https://bitbucket.org/szzoli/ite/.