Sharpening Jensen's Inequality
J. G. Liao, Arthur Berg
Introduction
Jensen’s inequality is a fundamental inequality in mathematics and it underlies many important statistical proofs and concepts. Some standard applications include derivation of the arithmetic-geometric mean inequality, non-negativity of Kullback and Leibler divergence, and the convergence property of the expectation-maximization algorithm (Dempster et al., 1977). Jensen’s inequality is covered in all major statistical textbooks such as Casella and Berger (2002, Section 4.7) and Wasserman (2013, Section 4.2) as a basic mathematical tool for statistics.
Let be a random variable with finite expectation and let be a convex function, then Jensen’s inequality (Jensen, 1906) establishes
We have incorporated the materials in this paper in our classroom teaching. With only slightly increased technical level and lecture time, we are able to present a much sharper version of the Jensen’s inequality that significantly enhances students’ understanding of the underlying concepts.
Main result
Let be a one-dimensional random variable with mean , and , where . Let is a twice differentiable function on , and define function
Let be the cumulative distribution function of . Applying Taylor’s theorem to about with a mean-value form of the remainder gives
where is between and . Explicitly solving for gives as defined above. Therefore
and the result follows because . ∎
Theorem 1 also holds when is replaced by and replaced by since
These less tight bounds are implied in the economics working paper Becker (2012). Our lower and upper bounds have the general form , where depends on . Similar forms of bounds are presented in Abramovich and Persson (2016); Dragomir (2014); Walker (2014), but our in Theorem 1 is much simpler and applies to a wider class of .
Inequality (2) implies Jensen’s inequality when . Note also that Jensen’s inequality is sharp when is linear, whereas inequality (2) is sharp when is a quadratic function of .
In some applications the moments of present in (2) are unknown, although a random sample from the underlying distribution is available. A version of Theorem 1 suitable for this situation is given in the following corollary.
Let be any datapoints in , and let
where and .
Consider the discrete random variable with probability distribution , . We have , , and . Then the corollary follows from application of Theorem 1. ∎
If is convex, then is monotonically increasing in , and if is concave, then is monotonically decreasing in .
We prove that when is convex. The analogous result for concave follows similarly. Note that
Without loss of generality we assume . Convexity of gives
Lemma 1 makes Theorem 1 easy to use as the follow results hold:
Note the limits of can be either finite or infinite. The proof of Lemma 1 borrows ideas from Bennish (2003). Examples of functions for which is convex include and for or . Examples of functions for which is concave include and for or .
Examples
For any random variable supported on with a finite variance, we can bound the moment generating function using Theorem 1 to get
For and , we have
So Theorem 1 provides no improvement over Jensen’s inequality. However, on a finite domain such as a non-negative random variable with , a significant improvement in the lower bound is possible because
The less sharp lower bound using is 0.125. Utilizing elaborate approximations and numerical optimizations Walker (2014) yielded a more accurate lower bound of 0.271.
Let be a positive random variable on interval with mean . Note that is convex whose derivative is concave. Applying Theorem 1 and Lemma 1 leads to
Now consider a sample of positive data points . Let be the arithmetic mean and be the geometric mean. Applying Corollary 1.1 gives
where , , are as defined in Corollary 1.1. To give some numerical results, we generated 100 random numbers from uniform distribution on . For these 100 numbers, the arithmetic mean is 54.830 and the geometric mean is 47.509. The above inequality becomes
which are fairly tight bounds. Replacing by and by leads to a less accurate lower bound 1.0339 and upper bound 21.698.
Let be a positive random variable on a positive interval with mean . For any real number , define the power mean as
Jensen’s inequality establishes that is an increasing function of . We now give a sharper inequality by applying Theorem 1. Let , , , and . Note that . Applying Theorem 1 leads to
To apply Lemma 1, note that is convex for or and is concave for or as noted in Section 2.
Applying the above result to the case of and , we have , . Therefore
For the same sequence generated in Example 2, we have . Applying Corollary 1.1 leads to
Note that the upper bound 48.905 is much smaller than the arithmetic mean by the Jensen’s inequality. Replacing by and by leads to a less accurate lower bound 0.8298 and 51.0839.
In a recent article published in the American Statistician, de Carvalho (2016) revisited Kolmogorov’s formulation of generalized mean as
where is a continuous monotone function with inverse . The Example 2 corresponds to and Example 3 corresponds to . We can also apply Theorem 1 to bound for a more general function .
Rao-Blackwell theorem (Theorem 7.3.17 in Casella and Berger, 2002; Theorem 10.42 in Wasserman, 2013) is a basic result in statistical estimation. Let be an estimator of , be a loss function convex in , and a sufficient statistic. Then the Rao-Blackwell estimator, , satisifies the following inequality in risk function
We can improve this inequality by applying Theorem 1 to with respect to the conditional distribution of given :
where function is defined as in Theorem 1 for and . Further taking expectations over gives
In particular for square-error loss, , we have
Using the original Jensen’s inequality only establishes the cruder inequality in Equation (4).
Improved bounds by partitioning
As discussed in Example 1 above, Theorem 1 does not improve on Jensen’s inequality if . In such cases, we can often sharpen the bounds by partitioning the domain following an approach used in Walker (2014). Let
, , and . It follows from the law of total expectation that
Let be a discrete random variable with distribution . It is easy to see that . It follows by Theorem 1 that
We can also apply Theorem 1 to each term:
Combining the above two equations, we have
Replacing by in the righthand side gives the upper bound.
The Jensen gap on the left side of (5) is positive if any of the terms on the right is positive. In particular, the Jensen gap is positive if there exists an interval that satisfies , and . Note that a finer partition does not necessarily lead to a sharper lower bound in (5). The focus of the partition should therefore be on isolating the part of interval in which is close to 0.
Consider example with and and . We divide into three intervals with equal probabilities. This gives
The actual Jensen gap is . The lower bound from (5) is 0.409, which is a huge improvement over Jensen’s bound of 0. The upper bound , however, provides no improvement over Theorem 1.
To summarize, this paper proposes a new sharpened version of the Jensen’s inequality. The proposed bound is simple and insightful, is broadly applicable by imposing minimum assumptions on , and provides fairly accurate result in spite of its simple form. It can be incorporated in any calculus-based statistical course.