Threshold Saturation on BMS Channels via Spatial Coupling
Shrinivas Kudekar, Cyril Measson, Tom Richardson, Ruediger Urbanke
I Introduction
It has long been known that convolutional LDPC ensembles, introduced by Felström and Zigangirov , have excellent thresholds when transmitting over general binary-input symmetric-output memoryless (BMS) channels. The fundamental reason underlying this good performance was recently discussed in detail in for the case when transmission takes place over the binary erasure channel (BEC).
In particular, it was shown in that the BP threshold of the spatially coupled ensemble is essentially equal to the MAP threshold of the underlying component ensemble. It was also shown that for long chains the MAP performance of the chain cannot be substantially larger than the MAP threshold of the component ensemble. In this sense, the BP threshold of the chain is increased to its maximal possible value. This is the reason why we call this phenomena threshold saturation via spatial coupling. In a recent paper , Lentmaier and Fettweis independently formulated the same statement as conjecture. They attribute the observation of the equality of the two thresholds to G. Liva.
It is tempting to conjecture that the same phenomenon occurs for transmission over general BMS channels. We provide some empirical evidence that this is indeed the case. In particular, we compute EBP GEXIT curves for transmission over the binary additive white Gaussian noise (BAWGN) channel. We show that these curves behave in an identical fashion to the ones when transmission takes place over the BEC. We also compute fixed points (FPs) of the spatial configuration and we demonstrate again empirically that these FPs have properties identical to the ones in the BEC case.
For a review on the literature on convolutional LDPC ensembles we refer the reader to and the references therein. As discussed in , there are many basic variants of coupled ensembles. For the sake of convenience of the reader, we quickly review the ensemble . This is the ensemble we use throughout the paper.
A discussion on the above ensemble and a proof of the following lemma can be found in .
The design rate of the ensemble , with , is given by
II Review: EBP GEXIT Curves, the Area Theorem, and the Maxwell Construction
Our aim is to empirically demonstrate that the performance of coupled ensembles is closely related to that of the underlying ensemble also in the general case. We limit our discussion to the coupling of regular ensembles. To get started, let us briefly review how the BP and MAP threshold can be characterized for regular ensembles. A detailed discussion can be found in .
For the BEC it is known that the behavior of the BP as well as the behavior of the MAP decoder are determined by the collection of FPs of DE. The same is conjectured to be true for general channels. Let us discuss this general conjecture.
The key concept in this conjecture is the EBP GEXIT curve . This curve is shown in Figure 1 for the -regular ensemble assuming that transmission takes place over the BAWGN. Numerically it is constructed in the following way. To construct one point on this curve find a FP , where denotes an element of the channel family under consideration. E.g., in the case considered in Figure 1, represents the -density of a BAWGN channel of variance . This FP gives rise to the point in the GEXIT plot. Hereby,
In words, computes the entropy associated to an -density, whereas computes the so-called GEXIT value of an -density. This GEXIT value depends on the “operating point”, i.e., it depends on the underlying channel . To first order, the GEXIT value is equal to the entropy.
We get the EBP GEXIT curve if we plot the points corresponding to all FPs of DE. For a detailed discussion we refer the reader to .
It was shown in that, under suitable technical conditions, for every GEXIT value there exists at least one FP with GEXIT value . Further, a simple recursive numerical procedure can be used to find such a FP. The technical difficulty lies in establishing the existence of the curve (rather than the existence of just the set of FPs). Although the numerical evidence strongly suggests the existence of the curve, it is an open problem to prove this analytically.
We can construct an upper bound on the MAP threshold as shown in Figure 2, see : integrate the EBP GEXIT curve starting from the point from right to left until the area under the curve equals the rate of the code. The point on the -axis where this equality occurs is an upper bound on the MAP threshold. It is conjectured to be in fact the exact MAP threshold.
III Density Evolution, Fixed Points, and the EBP GEXIT Curve for Coupled Ensembles
Let us describe the DE equations for the ensemble. In the sequel, densities are -densities. Let denote the channel density and let denote the density which is emitted by variable nodes at position .
Note that , where the r-h-s represents DE (without the effect of the channel) for the underlying -regular ensemble.
We write if is degraded w.r.t. . It is not hard to see that the function is monotone w.r.t. degradation in all its arguments , . More precisely, if we degrade any of the densities , , then is also degraded w.r.t. to its original value. We say that is monotone in its arguments. ∎
III-B Fixed Points and Admissible Schedules
Consider DE for the ensemble. Let . We call the constellation (of symmetric -densities). We say that forms a FP of DE with channel if fulfills (1) for . As a short hand we then say that is a FP. We say that is a non-trivial FP if is not identically equal to . Again, for , . ∎
III-C The EBP GEXIT Curve for Coupled Ensembles
We come now to the key point, the computation of the EBP GEXIT curve. From previous sections we have seen that FPs of forward DE are well defined and can be computed by applying a parallel schedule. This procedure allows us to compute stable FPs. As discussed in Section II, it was shown in how to compute unstable FPs for uncoupled ensembles by a modified DE procedure in which the entropy is kept fixed and the channel parameter is varied. The same procedure can be applied for coupled ensembles.
Figure 3 shows the result of this numerical computation when transmission takes place over the BAWGNC. Note that the resulting curves look very similar to the curves when transmission takes place over the BEC, see . For very small values of the curves are far to the right due to the significant rate loss that is incurred at the boundary. For around and above, the BP threshold of each ensemble is very close to the MAP threshold of the underlying -regular ensemble, namely . The picture strongly suggests that the same threshold saturation effect occurs for general channels as it was shown analytically to hold for the BEC in .
IV A Possible Proof Approach
So far the current discussion was empirical. Let us now quickly review which parts of the proof in can be extended easily and which currently seem difficult. (i) Constellation parameter: For the BEC, entropy is equal to the Bhattacharyya parameter which is equal to the erasure probability and which is also equal to twice the error probability. So, in this case any of those values is a natural quantity to parametrize constellations. For general channels, these parameters differ, and their choice is not necessarily equivalent for the purpose of the proof. (ii) Existence of FP: Although we did not explicitly state it in this short paper, the Existence Theorem 29 of , which guarantees the existence of a special FP of DE (c.f. Figure 4) can be extended to the general case by considering the Battacharyya functional. The main technical difficulty in the proof arises due to the fact that we are now operating on a space of symmetric probability densities. So to extend the proof of the BEC, we need to define appropriate metrics in this space so that we can apply the necessary fixed-point theorems. Together with the use of extremes of information combining methods, see , the proof extends. (iii) Shape of the constellation and the transition length: A key ingredient in the proof for the BEC was to show that any FP of DE has a very particular “shape.” More precisely, any FP had a very “fast” transition between its extreme values. Empirically, for the general case we observe the same phenomena. From Figure 4 we see that the FP quickly saturates to its maximum value (w.r.t. physical degradation) of the stable FP of the -regular ensemble. To show this property analytically seems currently to be one of the key difficulties in extending the proof. (iv) Construction of GEXIT Curve and the Area Theorem: Another key part of the BEC proof is the construction of a family of FPs (not necessarily stable FPs). The GEXIT curve plus the fast transition makes it possible in the case of the BEC to show that the “special” FP which was constructed via the existence theorem, has an associated channel parameter very close to the MAP threshold. How to best construct the GEXIT curve is an open issue.
V How to Mitigate the Rate Loss
We have seen that by coupling ensembles we can increase the threshold substantially. We also know that, due to the boundary condition, we incur a rate loss (see Lemma 1). This rate loss decreases linearly in , the length of the chain. Therefore, by picking large we can ensure that the rate loss is as small as desired. But large implies large codeword lengths and also may require a large number of iterations in the decoding process. This motivates to find ways of mitigating this rate loss.
To keep things simple, we consider transmission over the BEC. The same techniques and trade-offs apply to general BMS channels, but of course the given numerical values will change. We discuss two basic techniques: (i) rather than setting all boundary variables to be known, it suffices to set to known only a smaller fraction; (ii) it suffices to start the process at only one boundary rather than both. In addition there might be a benefit to consider ensembles defined on a circle rather than a line. This symmetrizes all positions of the ensemble, which in turn might lead to more efficient implementations.
Consider a “circular” ensemble. This ensemble is defined in a similar manner as the ensemble except that the positions are now from to and the index arithmetic is performed modulo . This circular ensemble has design rate equal to . If we let and if we set consecutive variable positions to then we recover the ensemble . This in itself gives a possibly more efficient way of implementing coupled ensembles. In this implementation all positions are symmetric, except for the received values, which are modified for the chosen consecutive positions.
Let us now generalize the construction. For let denote the fraction of variable nodes at position which we set to be known. E.g., if we set and all other values to then we recover the previous case. Define . As we will see shortly, from a rate perspective, it is not necessarily the best to set all variables at a certain position to . Further, it can be better to choose the “boundary” positions to be non-consecutive. This is why it is useful to introduce the above general model.
To start, let us compute the design rate for the above set-up. We denote this ensemble by .
The design rate of the ensemble is given by
The design rate is equal to , where and are the number check and variable nodes which are not fixed a priori to . Let us call those variable nodes “free.”
Let us start with . There are variable nodes per section. A fraction of those is permanently fixed to . Therefore, the number of free variable nodes is .
At each section there are check nodes. A check node imposes a constraint on the free variable nodes if at least one of its connections goes to a free variable node. The probability that a particular edge of a check node at position is connected to a frozen bit is equal to , where all index arithmetic is modulo . This implies that the probability that all edges of a check node at position are connected to frozen variables is equal to . Therefore, . ∎
Consider the ensemble where we set . We set all other values equal to . For the choice we know that we can achieve a threshold of roughly irrespective of the length of . What happens if we pick strictly less than ?
Let denote the “effective” erasure probability at the boundary. I.e., denotes the fraction of variables in the boundary which are free and which were erased during the transmission process. We have . What is the threshold that can be achieved (for arbitrary large ) for a given value of ?
Figure 5 shows the plot of according to DE for and as a function of . For e.g. , up to the achievable threshold is still equal to its maximal value, namely . For higher values of the threshold gracefully decreases. For we have . For this value of the corresponding rate for is . This is considerably larger than , which is the rate if . Indeed, the rate loss has almost been halved.
We can do slightly better. Pick and . Note that these two positions are non-contiguous. For this choice we still get a threshold of . The corresponding rate for is , which is slightly higher than in the previous case.
V-B One-Sided Ensemble
An alternative scheme is to define the ensemble on a line but to employ different terminations at the two boundaries. To be precise. Let the variable nodes be positioned from to . Assume the usual case. At the ”right” boundary use the following termination scheme. At position there are check nodes (instead of the usual ). Any edge, which under the standard connection rules should connect to a check node at a position or larger is mapped to a check node socket at position . In expectation, exactly such sockets are needed so that all check nodes at position have degree exactly . Therefore, locally the right-hand side boundary behaves exactly like a ensemble and there is no rate loss associated with this boundary.
On the left-hand side we proceed as in the standard ensemble. This reduces the rate-loss by a factor 2 compared to the standard ensemble. E.g., for we get for our usual -ensemble a rate of . But we can do better.
Let and consider an ensemble in which the right boundary is terminated without rate loss as described above. Under the standard scheme, the check nodes at position have degree and check nodes at position have degree .
Take a fraction of the check nodes at position and merge them with check nodes at position . Those merged nodes at position have degree as in the regular case. As long as is sufficiently small the threshold still remains unchanged. But this merging further reduces the incurred rate loss.
A small note of caution might be in order at this place. Even though we can, as we just showed, mitigate the rate loss, this comes at some price – the number of required iterations will go up. A characterization of the involved trade off would be of high practical interest.
VI Conclusion
Starting with the work by Felström and Zigangirov , it has been known that coupled ensembles have an excellent performance. This was confirmed via threshold computations by Sridharan, Lentmaier, Costello and Zigangirov for the BEC and by Lentmaier, Sridharan, Zigangirov and Costello for general channels . In it was shown that for transmission over the BEC the BP behavior of coupled ensembles is essentially equal to the MAP behavior of the underlying ensemble.
The current paper provides numerical evidence which suggest that exactly the same behavior occurs also for transmission over general BMS channels. We have further extended some of the basic techniques and statements which were used in to accomplish the proof for the BEC to the general setup. A complete proof is unfortunately still open. As discussed already in , such a proof would automatically show that it is possible to achieve capacity under iterative coding, and in addition, the convergence to capacity would be uniform over the whole class of BMS channels.
VII Acknowledgments
SK acknowledges support of NMC via the NSF collaborative grant CCF-0829945 on “Harnessing Statistical Physics for Computing and Communications”. SK would also like to thank Misha Chertkov for his encouragement.