Deciphering the global organization of clustering in real complex networks

Pol Colomer-de-Simon, M. Angeles Serrano, Mariano G. Beiro, J. Ignacio Alvarez-Hamelin, Marian Boguna

I Introduction

The architecture of real complex systems lay at the midpoint between order and disorder, although its precise location is quite difficult to determine. Disorder in complex networks is manifested by the small-world effect Watts and Strogatz 1998 and a highly heterogeneous degree distribution Barabási and Albert 1999, both properties commonly present in real complex networks Dorogovtsev and Mendes 2003; Newman 2010. Order is, on the other hand, manifested by the presence of triangles –or clustering– representing three point correlations in the system. Indeed, the very concept of order is typically related to the existence of a metric structure in the system which, from the network perspective, is captured by clustering, the smallest network motif able to encode the triangle inequality. Yet, unlike the small-world effect and the heterogeneity of nodes’ degrees, clustering is not an emergent property spontaneously generated by paradigmatic connectivity principles such as preferential attachment and, therefore, calls for specific mechanisms for explaining its emergence, thus giving important insights into the nature of network formation and network evolution.

On the other hand, the effects of clustering on the structural and dynamical properties of networks have not yet been conclusively elucidated. In fact, several studies have reported apparently contradictory results concerning the effects of clustering on the percolation properties of networks and little is known on its effects on dynamical processes running on networks Serrano and Boguñá 2006a; Serrano and Boguñá 2006b; Trapman 2007; Newman 2009; Gleeson 2009; Gleeson and Melnik 2009; Gleeson et al. 2010. This is further hindered by the technical difficulties of any analytical treatment. Indeed, the presence of strong clustering invalidates, in general, the “locally tree-like” assumption used in random graphs, leaving little room for any theoretical study. In an effort to overcome these problems, a new class of clustered network models has been proposed. These models start by defining a certain set of cliques (fully connected subgraphs) of different sizes that are afterwords connected in a random fashion. In this way, by considering cliques as super-nodes, the network connecting these super-nodes is locally tree-like, thus allowing for an analytical treatment Trapman 2007; Newman 2009; Gleeson 2009; Gleeson and Melnik 2009; Karrer and Newman 2010; Gleeson et al. 2010; Allard et al. 2012. Then, it is possible to generate networks with a given degree distribution P(k)P(k) and degree-dependent clustering coefficient cˉ(k)\bar{c}(k), defined as the average fraction of triangles attached to nodes of degree kk.

While this is indeed a fair approach to the problem, triangles generated by these models are arranged in a very specific way, with strong correlations between the properties of adjacent edges. In some sense, we can consider this class of models as generators of maximally ordered clustered graphs. At the other side of the spectrum, we can define an ensemble of maximally random clustered graphs such that correlations among adjacent edges are the minimum needed to conform with the degree-dependent clustering coefficient, but no more. These two types of models define –in a non-rigorous way– two extremes of the phase space of possible graphs with given P(k)P(k) and cˉ(k)\bar{c}(k). A simple question arises then: where are real networks positioned in this phase space? To give an answer to this question, we need to go beyond the local properties of networks and to study their global organization. In this paper, we study the global structure of clustering in real networks and compare them with the global structure of clustering induced by the two types of models with identical local properties. More specifically, we analyze the organization of real and model networks into mm-cores, defined as maximal subgraphs with edges participating at least in mm triangles, that is able to distinguish between hierarchical and modular architectures. Interestingly enough, real networks tend to be closer to maximally random clustered graphs, although clear differences are evident.

II Results

In this paper, we analyze three real paradigmatic networks from different domains: the Internet at the Autonomous System level Boguñá et al. 2010, the web of trust of the Pretty Good Privacy protocol (PGP) Boguñá et al. 2004a, and the metabolic network of the bacterium E. coli Serrano et al. 2012. However, the results obtained here also hold for a wide spectrum of systems (See Supplementary Information for the analysis of a larger set of systems). We first describe their random counterparts, namely, maximally ordered and maximally random clustered graphs with the same degree distribution and clustering spectrum.

One of the best clique-based models to generate maximally ordered clustered networks is the one introduced by Gleeson in Gleeson 2009. In this model, nodes belong to single cliques and are also given a number of connections outside their cliques. Then, cliques are considered as super-nodes, each with an effective degree given by the sum of all the external links of the members of the clique, and connected using the standard configuration model. The input of the model is the joint distribution γ(c,k)\gamma(c,k), defined as the probability that a randomly chosen node has degree kk and belongs to a clique of size cc. Both the degree distribution and the degree-dependent clustering coefficient are related to function γ(c,k)\gamma(c,k). Therefore, by properly choosing its form, it is possible to match the desired degree distribution and clustering. Note, however, that since we start with cliques and not nodes, the number of nodes and their actual degrees are not fixed a priori. As a consequence, in finite heterogeneous networks, there may be some unavoidable discrepancies between real and random versions of the network. Hereinafter, we denote this model as “clique-based model” (CB).

On the other hand, we generate maximally random clustered networks as an ensemble of exponential graphs Park and Newman 2004 with Hamiltonian

where cˉ(k)\bar{c}(k) is the target degree-dependent clustering coefficient and cˉ∗(k)\bar{c}^{*}(k) is the one corresponding to the current state of the network. This Hamiltonian is minimized by means of simulated annealing coupled to a Metropolis rewiring scheme until the current clustering is close enough to the target one (see Methods Section for further details). Here we use two different rewiring schemes. In the first one Maslov and Sneppen 2002, degrees of nodes are preserved after each single rewiring event but correlations between the degrees of connected nodes are either destroyed or brought down to the level of the structural ones Boguñá et al. 2004b; Serrano et al. 2007. In the second scheme Melnik et al. 2011, rewiring events preserve both the degree distribution and the joint degree-degree distribution of connected nodes, P(k,k′)P(k,k^{\prime}), so that degree-degree correlations are fully preserved. Hereinafter, we denote these models as “maximally random models” (MR). We would like to stress that, even though there are many models of exponential random graphs generating clustered graphs Frank and Strauss 1986; Milo et al. 2002; Foster et al. 2010, none of them reproduces the actual clustering spectrum as a function of node degree. In this sense, our maximally random model gets closer to real networks.

Notice that none of the random models used in this paper enforces global connectivity of the network in a single connected component. Therefore, the number of disconnected components and the size of the giant (or largest) component must be considered as predictions of the models, which can be readily compared to those of real networks. In Table 1, we show this comparison with the networks analyzed in this paper. Quite remarkably, in the case of the Internet, MR models predict the existence of, basically, a single connected component, as it is also observed in the real network. On the other hand, the CB model generates a very large number of disconnected components and a giant component significantly smaller than the real one. Even more surprising are the results for the PGP web of trust. The real network is fragmented into a large number of small components whereas its giant component occupies around 18% of the network. All models generate a similar number of disconnected components. However, the relative size of the giant component is very well reproduced by MR models, whereas the CB model predicts a giant component twice as large. In the case of the metabolic network of the bacterium E. coli, all models predicts the existence of a single connected component, in good agreement with the real network.

II.2 Revealing network hierarchies: kk-cores and mm-cores

Real heterogeneous networks are typically hierarchically organized. One of the most useful tools to uncover such hierarchies is the kk-core decomposition Dorogovtsev et al. 2006. Given a network, its kk-core is defined as the maximal subgraph such that all nodes in the subgraph have at least kk connections with members of the subgraph. This defines a hierarchy of nested subgraphs, where the 11-core contains the 22-core, which in turn contains the 33-core and so on until the maximum kk-core is reached. Nodes belonging to the kk-core but not to the (k+1)(k+1)-core are said to have coreness kk. Real networks often show a deep and complex kk-core structure, as made evident by tools such as LaNet-vi Beiró et al. 2008. However, even though clustering has been shown to induce strong kk-core hierarchiesSerrano and Boguñá 2006a, the kk-core per se does not include any information about clustering and, thus, cannot discriminate well between two networks with different global organization of clustering but with the same clustering coefficient.

To overcome this problem, the concept of kk-core has been remodeled to account for clustered networks. A key ingredient throughout the paper is the concept of edge multiplicity mm, defined as the number of distinct triangles going through an edge Radicchi et al. 2004; Serrano and Boguñá 2006c. All edges belonging to a clique of size nn have identical multiplicity n−2n-2 whereas an edge connecting two cliques has zero multiplicity. Therefore, strong correlations between the multiplicities of adjacent edges indicate that triangles are arranged in a clique-like fashion whereas a weaker correlation indicate a random distribution of triangles. It is therefore clear that, in order to uncover the global organization of triangles in a network, it is necessary to understand the organization of the multiplicities of their edges. This can be achieved with the mm-core, defined as the maximal subgraph such that all its edges have, at least, multiplicity mm within it. This concept was developed in Saito et al. 2008; Gregori et al. 2013 under the name of kk-dense decomposition. The edges in a kk-dense graph have multiplicity m=k−2m=k-2. Because of this, we prefer the notion of mm-core, which is directly related to the multiplicity: an edge belongs to the mm-core if its multiplicity within the mm-core is, at least, mm. A node belongs to the mm-core if at least one of its edges belongs to it. A node belonging to the mm-core but not to the (m+1)(m+1)-core is said to have mm-coreness mm. As in the case of the kk-core, the mm-core defines a set of nested subgraphs whose properties informs us about the global organization of triangles in the graph. The left plot in Fig. 1 shows an example of a simple network and its mm-core structure.

In the case of the kk-core, the density of links within each subgraph grows as kk is increased. As a consequence, it is very unlikely that the (k+1)(k+1)-core is fragmented in different components if the kk-core is connected. Therefore, the main interest of the kk-core decomposition is focused on the size of the giant kk-core and the maximum coreness of the system. The situation is completely different in the case of the mm-core. This is so because of a weaker correlation between mm-coreness of a node and its degree Orsini et al. 2013. In fact, the mm-core decomposition is able to distinguish between a strong hierarchical structure –when mm-cores do not fragment into smaller components– from a highly modular architecture –when mm-cores are always fragmented. In this case, the quantities of interest are, besides the size of the giant mm-core and the maximum mm-coreness, the number of components as a function of mm.

Figures 2, 4, and 6 show a comparison of the kk-core and mm-core decompositions between real networks and their random equivalents. As it can be observed in the top plots of these figures, all models do a reasonably good job at reproducing both the kk-core structure and the distribution of edge multiplicities, even though MR models are clearly better than the CB one. However, there are important differences in the mm-core decomposition. While both versions of MR models reproduce well the giant mm-core, the maximum mm-coreness, and the number of components as a function of mm of all the studied networks, the CB model overestimates the size and number of components in the case of the Internet and underestimate the size of giant mm-cores in the PGP web of trust. In the case of the metabolic network, MR models reproduce well its entire mm-core structure. The CB model, on the other hand, does not capture well the mm-core decomposition. Even though the CB network is originally connected, it fragments into a large number of disconnected components already at the m1m1-core and keeps fragmenting at each level almost up to the largest mm-core, which is also three times larger than the real one.

II.3 mm-core visualization

The mm-core decomposition is actually much richer and complex than what Figs. 2, 4, and 6 show. Certainly, the mm-core decomposition can be represented as a branching process that encodes the fragmentation of mm-cores into disconnected components as mm is increased. The tree-like structure of this process informs us about the global organization –for instance hierarchical vs. modular– of clustering in networks. To visualize this process we use LaNet-vi 3.0 Alvarez-Hamelin et al. 2013, a modified version of LaNet-vi Beiró et al. 2008, originally designed to visualize the kk-core structure of a network. In short, LaNet-vi tool evaluates the coreness of all nodes of the network and arrange them in a plane following the hierarchy induced by the kk-cores, so that nodes with high coreness are placed at the center of the figure whereas nodes with lower coreness are located around nodes with higher coreness in an onion-like shape. The major modification in LaNet-vi 3.0 with respect to the visualization mode in the previous version concerns the representation of disconnected components. If the network forms a single connected component, nodes with mm-coreness 0 are arranged in the outermost circle of the representation. Whenever the m1m1-core is fragmented into several components, these are arranged in separate and non-overlapping disks within the circle of mm-coreness 0, with nodes of mm-coreness 1 placed at the edge of their corresponding disk. The process is repeated for each disconnected component with the m2m2-core, m3m3-core, etc., until the maximum mm-coreness present in the network is reached. The size of each disk is proportional to the logarithm of the number of nodes in the component. In this way, it is possible to visualize simultaneously all the information encoded in the mm-cores so that different networks can be easily compared (see the right plot in Fig. 1 for a simple example). When the original network is already fragmented (like in the PGP web of trust, for instance), we first proceed to arrange disconnected components in non overlapping disks within the outermost disk, that in this case does not have any node in its perimeter.

Figures 3, 5, and 7 show the visualization of mm-cores of real networks and their random equivalents (visualizations of MR models are shown only for P(k)P(k) preserving rewiring). In the case of the Internet graph, the mm-core visualization reveals a strongly hierarchical structure, where each layer is contained within the previous layer and where connections are mainly radial, with nodes with low mm-coreness connected to nodes with higher mm-coreness and very few connections between nodes in the same layer. Interestingly, this type of structure is also revealed in recent embeddings of the Internet graph into the hyperbolic plane Boguñá et al. 2010. This structure is very well reproduced by MR models, as it can be seen in the left bottom plot of Fig. 3, but not by the CB model, which generates a highly modular and non-hierarchical structure. The case of the web of trust of PGP is particularly interesting. Figure 5 reveals a mixture of a modular structure, with a strong fragmentation for all values of mm –as one would expect for a social network–, and a hierarchical structure, revealed by the existence of a persistent giant mm-core and a large number of layers. Again, this structure is very well reproduced by MR models whereas the CB model generates a very flat modular structure without any hierarchy. Finally, the metabolic network is also strongly hierarchical, although due to the small network size the number of layers is relatively small. MR models reproduce very well its structure whereas the CB model does not generate any hierarchy.

III Discussion

The results presented in this paper indicate that, in agreement with previous studies Jamakovic et al. 2009; Foster et al. 2011, the degree distribution P(k)P(k) and clustering spectrum cˉ(k)\bar{c}(k) are the main contributors to the global organization of the majority of real networks, which are close to maximally random once these properties are fixed. This supports the idea that most real networks are the result of a self-organized process based on local optimization rules, in contrast to global optimization principles, that yield a hierarchical organization that cannot be reproduced by maximally ordered clustered models. Besides, the strong clustering observed in real networks, supports also the idea that such local principles are related to a similarity measure among nodes of the network that can be quantified by an underlying metric structure Serrano et al. 2008; Boguñá et al. 2009; Boguñá et al. 2010; Krioukov et al. 2010; Serrano et al. 2012; Papadopoulos et al. 2012. On the other hand, global optimization principles are necessarily present, for instance, in power grids, where they induce topologies that are very different from what one would expect at random. This is made evident by its mm-core decomposition (see Supplementary Information). In this case, even thought the mm-core structure is not very deep, it is very different from any of the random models, which generate highly unstructured mm-cores. Therefore, the mm-core decomposition along with its visualization tool can help us to find the true mechanisms at play in the formation and evolution of real networks.

IV Methods

Maximally random clustered networks are generated by means of a biased rewiring procedure. We use two different rewiring schemes. In the first one, two different edges are chosen at random. Let these connect nodes A with B and C with D. Then, the two edges are swapped so that nodes A and D, on the one hand, and C and B, on the other, are now connected. We take care that no self-connections or multiple connection between the same pair of nodes are induced by this process. This rewiring scheme preserves the degree distribution of the original network but not degree-degree correlations. In the second rewiring scheme, we first chose an edge at random and look at the degree of one of its attached nodes, kk. Then, a second link attached to a node of the same degree kk is chosen and the two links are swapped as before. Notice that this procedure preserves both the degree of each node and the actual nodes’ degrees at the end of the two original edges. Therefore, the procedure preserves the full degree-degree correlation structure encoded in the joint distribution P(k,k′)P(k,k^{\prime}). Both procedures are ergodic and satisfy detailed balance.

Regardless of the rewiring scheme at use, the process is biased so that generated graphs belong to an exponential ensemble of graphs G={G}\cal{G}=\{G\}, where each graph has a sampling probability P(G)∝e−βH(G)P(G)\propto e^{-\beta H(G)}, where β\beta is the inverse of the temperature and H(G)H(G) is a Hamiltonian that depends on the current network configuration. Here we consider ensembles where the Hamiltonian depends on the target clustering spectrum of the real Network cˉ(k)\bar{c}(k) as

where cˉ∗(k)\bar{c}^{*}(k) is the current degree-dependent clustering coefficient. We then use a simulated annealing algorithm based on a standard Metropolis-Hastings procedure. Let G′G^{\prime} be the new graph obtained after one rewiring event, as defined above. The candidate network G′G^{\prime} is accepted with probability

otherwise, we keep the graph GG unchanged. We first start by rewiring the real network 200E200E times at β=0\beta=0, where EE is the total number of edges of the network. This step destroys the clustering coefficient of the original network. Then, we start an annealing procedure at β0=50\beta_{0}=50, increasing the parameter β\beta by a 10%10\% after 100E100E rewiring events have taken place. We keep increasing β\beta until the target clustering spectrum is reached within a predefined precision or no further improvement can be achieved.

IV.2 Computing mm-cores

To compute mm-cores efficiently, we develop a new approach, different from the one in Saito et al. 2008; Gregori et al. 2013. We first map the original graph GG into a hypergraph G∗G^{*}, where edges in GG become vertices in G∗G^{*} and where each triangle in the original graph is mapped into an edge (a 33-tuple) in G∗G^{*}. Then, by noticing that the degree of a vertex v∗v^{*} in G∗G^{*} equals the number of triangles associated to the original edge in GG, it is possible to obtain the mm-core just by computing the kk-core of the same level in G∗G^{*}. The complete description can be found in the Supplementary Information.

References