On Relations Between the Relative Entropy and χ2-Divergence, Generalizations and Applications

Nishiyama, Tomohiro; Sason, Igal

doi:10.3390/e22050563

Open AccessArticle

On Relations Between the Relative Entropy and χ²-Divergence, Generalizations and Applications

by

Tomohiro Nishiyama

¹

and

Igal Sason

^2,*

¹

Independent Researcher, Tokyo 206–0003, Japan

²

Faculty of Electrical Engineering, Technion—Israel Institute of Technology, Technion City, Haifa 3200003, Israel

^*

Author to whom correspondence should be addressed.

Entropy 2020, 22(5), 563; https://doi.org/10.3390/e22050563

Submission received: 22 April 2020 / Revised: 12 May 2020 / Accepted: 17 May 2020 / Published: 18 May 2020

(This article belongs to the Special Issue Divergence Measures: Mathematical Foundations and Applications in Information-Theoretic and Statistical Problems)

Download Versions Notes

Abstract

:

This paper is focused on a study of integral relations between the relative entropy and the chi-squared divergence, which are two fundamental divergence measures in information theory and statistics, a study of the implications of these relations, their information-theoretic applications, and some generalizations pertaining to the rich class of f-divergences. Applications that are studied in this paper refer to lossless compression, the method of types and large deviations, strong data–processing inequalities, bounds on contraction coefficients and maximal correlation, and the convergence rate to stationarity of a type of discrete-time Markov chains.

Keywords:

relative entropy; chi-squared divergence; f-divergences; method of types; large deviations; strong data–processing inequalities; information contraction; maximal correlation; Markov chains

1. Introduction

The relative entropy (also known as the Kullback–Leibler divergence [1]) and the chi-squared divergence [2] are divergence measures which play a key role in information theory, statistics, learning, signal processing, and other theoretical and applied branches of mathematics. These divergence measures are fundamental in problems pertaining to source and channel coding, combinatorics and large deviations theory, goodness-of-fit and independence tests in statistics, expectation–maximization iterative algorithms for estimating a distribution from an incomplete data, and other sorts of problems (the reader is referred to the tutorial paper by Csiszár and Shields [3]). They both belong to an important class of divergence measures, defined by means of convex functions f, and named f-divergences [4,5,6,7,8]. In addition to the relative entropy and the chi-squared divergence, this class unifies other useful divergence measures such as the total variation distance in functional analysis, and it is also closely related to the Rényi divergence which generalizes the relative entropy [9,10]. In general, f-divergences (defined in Section 2) are attractive since they satisfy pleasing features such as the data–processing inequality, convexity, (semi)continuity, and duality properties, and they therefore find nice applications in information theory and statistics (see, e.g., [6,8,11,12]).

In this work, we study integral relations between the relative entropy and the chi-squared divergence, implications of these relations, and some of their information-theoretic applications. Some generalizations which apply to the class of f-divergences are also explored in detail. In this context, it should be noted that integral representations of general f-divergences, expressed as a function of either the DeGroot statistical information [13], the

E_{γ}

-divergence (a parametric sub-class of f-divergences, which generalizes the total variation distance [14] [p. 2314]) and the relative information spectrum, have been derived in [12] [Section 5], [15] [Section 7.B], and [16] [Section 3], respectively.

Applications in this paper are related to lossless source compression, large deviations by the method of types, and strong data–processing inequalities. The relevant background for each of these applications is provided to make the presentation self contained.

We next outline the paper contributions and the structure of our manuscript.

1.1. Paper Contributions

This work starts by introducing integral relations between the relative entropy and the chi-squared divergence, and some inequalities which relate these two divergences (see Theorem 1, its corollaries, and Proposition 1). It continues with a study of the implications and generalizations of these relations, pertaining to the rich class of f-divergences. One implication leads to a tight lower bound on the relative entropy between a pair of probability measures, expressed as a function of the means and variances under these measures (see Theorem 2). A second implication of Theorem 1 leads to an upper bound on a skew divergence (see Theorem 3 and Corollary 3). Due to the concavity of the Shannon entropy, let the concavity deficit of the entropy function be defined as the non-negative difference between the entropy of a convex combination of distributions and the convex combination of the entropies of these distributions. Then, Corollary 4 provides an upper bound on this deficit, expressed as a function of the pairwise relative entropies between all pairs of distributions. Theorem 4 provides a generalization of Theorem 1 to the class of f-divergences. It recursively constructs non-increasing sequences of f-divergences and as a consequence of Theorem 4 followed by the usage of polylogairthms, Corollary 5 provides a generalization of the useful integral relation in Theorem 1 between the relative entropy and the chi-squared divergence. Theorem 5 relates probabilities of sets to f-divergences, generalizing a known and useful result by Csiszár for the relative entropy. With respect to Theorem 1, the integral relation between the relative entropy and the chi-squared divergence has been independently derived in [17], which also derived an alternative upper bound on the concavity deficit of the entropy as a function of total variational distances (differing from the bound in Corollary 4, which depends on pairwise relative entropies). The interested reader is referred to [17], with a preprint of the extended version in [18], and to [19] where the connections in Theorem 1 were originally discovered in the quantum setting.

The second part of this work studies information-theoretic applications of the above results. These are ordered by starting from the relatively simple applications, and ending at the more complicated ones. The first one includes a bound on the redundancy of the Shannon code for universal lossless compression with discrete memoryless sources, used in conjunction with Theorem 3 (see Section 4.1). An application of Theorem 2 in the context of the method of types and large deviations analysis is then studied in Section 4.2, providing non-asymptotic bounds which lead to a closed-form expression as a function of the Lambert W-function (see Proposition 2). Strong data–processing inequalities with bounds on contraction coefficients of skew divergences are provided in Theorem 6, Corollary 7 and Proposition 3. Consequently, non-asymptotic bounds on the convergence to stationarity of time-homogeneous, irreducible, and reversible discrete-time Markov chains with finite state spaces are obtained by relying on our bounds on the contraction coefficients of skew divergences (see Theorem 7). The exact asymptotic convergence rate is also obtained in Corollary 8. Finally, a property of maximal correlations is obtained in Proposition 4 as an application of our starting point on the integral relation between the relative entropy and the chi-squared divergence.

1.2. Paper Organization

This paper is structured as follows. Section 2 presents notation and preliminary material which is necessary for, or otherwise related to, the exposition of this work. Section 3 refers to the developed relations between divergences, and Section 4 studies information-theoretic applications. Proofs of the results in Section 3 and Section 4 (except for short proofs) are deferred to Section 5.

2. Preliminaries and Notation

This section provides definitions of divergence measures which are used in this paper, and it also provides relevant notation.

Definition 1.

[12] [p. 4398] Let P and Q be probability measures, let μ be a dominating measure of P and Q (i.e.,

P, Q ≪ μ

), and let

p : = \frac{d P}{d μ}

and

q : = \frac{d Q}{d μ}

be the densities of P and Q with respect to μ. The f-divergence from P to Q is given by

\begin{matrix} D_{f} (P ∥ Q) : = \int q f (\frac{p}{q}) d μ, \end{matrix}

(1)

where

\begin{matrix} f (0) : = lim_{t \to 0^{+}} f (t), 0 f (\frac{0}{0}) : = 0, \end{matrix}

(2)

\begin{matrix} 0 f (\frac{a}{0}) : = lim_{t \to 0^{+}} t f (\frac{a}{t}) = a lim_{u \to \infty} \frac{f (u)}{u}, a > 0 . \end{matrix}

(3)

It should be noted that the right side of (1) does not depend on the dominating measure μ.

Throughout the paper, we denote by

1 {relation}

the indicator function; it is equal to 1 if the relation is true, and it is equal to 0 otherwise. Throughout the paper, unless indicated explicitly, logarithms have an arbitrary common base (that is larger than 1), and

exp (\cdot)

indicates the inverse function of the logarithm with that base.

Definition 2.

[1] The relative entropy is the f-divergence with

f (t) : = t log t

for

t > 0

,

\begin{matrix} (4) & D (P ∥ Q) & : = D_{f} (P ∥ Q) \\ (5) & = \int p log \frac{p}{q} d μ . \end{matrix}

Definition 3.

The total variation distance between probability measures P and Q is the f-divergence from P to Q with

f (t) : = | t - 1 |

for all

t \geq 0

. It is a symmetric f-divergence, denoted by

| P - Q |

, which is given by

\begin{matrix} (6) & | P - Q | & : = D_{f} (P ∥ Q) \\ (7) & = \int | p - q | d μ . \end{matrix}

Definition 4.

[2] The chi-squared divergence from P to Q is defined to be the f-divergence in (1) with

f (t) : = {(t - 1)}^{2}

or

f (t) : = t^{2} - 1

for all

t > 0

,

\begin{matrix} (8) & χ^{2} (P ∥ Q) & : = D_{f} (P ∥ Q) \\ (9) & = \int \frac{{(p - q)}^{2}}{q} d μ = \int \frac{p^{2}}{q} d μ - 1 . \end{matrix}

The Rényi divergence, a generalization of the relative entropy, was introduced by Rényi [10] in the special case of finite alphabets. Its general definition is given as follows (see, e.g., [9]).

Definition 5.

[10] Let P and Q be probability measures on

X

dominated by μ, and let their densities be respectively denoted by

p = \frac{d P}{d μ}

and

q = \frac{d Q}{d μ}

. The Rényi divergence of order

α \in [0, \infty]

is defined as follows:

If $α \in (0, 1) \cup (1, \infty)$ , then

$\begin{matrix} (10) & D_{α} (P ∥ Q) & = \frac{1}{α - 1} log E [p^{α} (Z) q^{1 - α} (Z)] \\ (11) & = \frac{1}{α - 1} log \sum_{x \in X} P^{α} (x) Q^{1 - α} (x), \end{matrix}$

where $Z \sim μ$ in (10), and (11) holds if $X$ is a discrete set.
By the continuous extension of $D_{α} (P ∥ Q)$ ,

$\begin{matrix} (12) & D_{0} (P ∥ Q) = max_{A : P (A) = 1} log \frac{1}{Q (A)}, \\ (13) & D_{1} (P ∥ Q) = D (P ∥ Q), \\ (14) & D_{\infty} (P ∥ Q) = log ess sup \frac{p (Z)}{q (Z)} . \end{matrix}$

The second-order Rényi divergence and the chi-squared divergence are related as follows:

\begin{matrix} D_{2} (P ∥ Q) = log (1 + χ^{2} (P ∥ Q)), \end{matrix}

(15)

and the relative entropy and the chi-squared divergence satisfy (see, e.g., [20] [Theorem 5])

\begin{matrix} D (P ∥ Q) \leq log (1 + χ^{2} (P ∥ Q)) . \end{matrix}

(16)

Inequality (16) readily follows from (13), (15), and since

D_{α} (P ∥ Q)

is monotonically increasing in

α \in (0, \infty)

(see [9] [Theorem 3]). A tightened version of (16), introducing an improved and locally-tight upper bound on

D (P ∥ Q)

as a function of

χ^{2} (P ∥ Q)

and

χ^{2} (Q ∥ P)

, is introduced in [15] [Theorem 20]. Another sharpened version of (16) is derived in [15] [Theorem 11] under the assumption of a bounded relative information. Furthermore, under the latter assumption, tight upper and lower bounds on the ratio

\frac{D (P ∥ Q)}{χ^{2} (P ∥ Q)}

are obtained in [15] [(169)].

Definition 6.

[21] The Györfi–Vajda divergence of order

s \in [0, 1]

is an f-divergence with

\begin{matrix} f (t) = ϕ_{s} (t) : = \frac{{(t - 1)}^{2}}{s + (1 - s) t}, t \geq 0 . \end{matrix}

(17)

Vincze–Le Cam distance (also known as the triangular discrimination) ([22,23]) is a special case with

s = \frac{1}{2}

.

In view of (1), (9) and (17), it can be verified that the Györfi–Vajda divergence is related to the chi-squared divergence as follows:

D_{ϕ_{s}} (P ∥ Q) = {\begin{cases} \frac{1}{s^{2}} \cdot χ^{2} (P ∥ (1 - s) P + s Q), & s \in (0, 1], \\ χ^{2} (Q ∥ P), & s = 0 . \end{cases}

(18)

Hence,

\begin{matrix} (19) & D_{ϕ_{1}} (P ∥ Q) = χ^{2} (P ∥ Q), \\ (20) & D_{ϕ_{0}} (P ∥ Q) = χ^{2} (Q ∥ P) . \end{matrix}

3. Relations between Divergences

We introduce in this section results on the relations between the relative entropy and the chi-squared divergence, their implications, and generalizations. Information–theoretic applications are studied in the next section.

3.1. Relations between the Relative Entropy and the Chi-Squared Divergence

The following result relates the relative entropy and the chi-squared divergence, which are two fundamental divergence measures in information theory and statistics. This result was recently obtained in an equivalent form in [17] [(12)] (it is noted that this identity was also independently derived by the coauthors in two separate un-published works in [24] [(16)] and [25]). It should be noted that these connections between divergences in the quantum setting were originally discovered in [19] [Theorem 6]. Beyond serving as an interesting relation between these two fundamental divergence measures, it is introduced here for the following reasons:

(a): New consequences and applications of it are obtained, including new shorter proofs of some known results;
(b): An interesting extension provides new relations between f-divergences (see Section 3.3).

Theorem 1.

Let P and Q be probability measures defined on a measurable space

(X, F)

, and let

\begin{matrix} R_{λ} : = (1 - λ) P + λ Q, λ \in [0, 1] \end{matrix}

(21)

be the convex combination of P and Q. Then, for all

λ \in [0, 1]

,

\begin{matrix} \frac{1}{log e} D (P ∥ R_{λ}) & = \int_{0}^{λ} χ^{2} (P ∥ R_{s}) \frac{d s}{s}, \end{matrix}

(22)

\begin{matrix} \frac{1}{2} λ^{2} χ^{2} (R_{1 - λ} ∥ Q) & = \int_{0}^{λ} χ^{2} (R_{1 - s} ∥ Q) \frac{d s}{s} . \end{matrix}

(23)

Proof.

See Section 5.1. □

A specialization of Theorem 1 by letting

λ = 1

gives the following identities.

Corollary 1.

\begin{matrix} \frac{1}{log e} D (P ∥ Q) = \int_{0}^{1} χ^{2} (P ∥ (1 - s) P + s Q) \frac{d s}{s}, \end{matrix}

(24)

\begin{matrix} \frac{1}{2} χ^{2} (P ∥ Q) = \int_{0}^{1} χ^{2} (s P + (1 - s) Q ∥ Q) \frac{d s}{s} . \end{matrix}

(25)

Remark 1.

The substitution

s : = \frac{1}{1 + t}

transforms (24) to [26] [Equation (31)], i.e.,

\begin{matrix} \frac{1}{log e} D (P ∥ Q) = \int_{0}^{\infty} χ^{2} (P ∥ \frac{t P + Q}{1 + t}) \frac{d t}{1 + t} . \end{matrix}

(26)

In view of (18) and (21), an equivalent form of (22) and (24) is given as follows:

Corollary 2.

For

s \in [0, 1]

, let

ϕ_{s} : [0, \infty) \to R

be given in (17). Then,

\begin{matrix} (27) & \frac{1}{log e} D (P ∥ R_{λ}) & = \int_{0}^{λ} s D_{ϕ_{s}} (P ∥ Q) d s, λ \in [0, 1], \\ (28) & \frac{1}{log e} D (P ∥ Q) & = \int_{0}^{1} s D_{ϕ_{s}} (P ∥ Q) d s . \end{matrix}

By Corollary 1, we obtain original and simple proofs of new and old f-divergence inequalities.

Proposition 1.

(f-divergence inequalities).

(a): Pinsker’s inequality:

$\begin{matrix} D (P ∥ Q) \geq \frac{1}{2} {| P - Q |}^{2} log e . \end{matrix}$

(29)
(b): $\begin{matrix} \frac{1}{log e} D (P ∥ Q) \leq \frac{1}{3} χ^{2} (P ∥ Q) + \frac{1}{6} χ^{2} (Q ∥ P) . \end{matrix}$

(30)

Furthermore, let ${P_{n}}$ be a sequence of probability measures that is defined on a measurable space $(X, F)$ , and which converges to a probability measure P in the sense that

$\begin{matrix} lim_{n \to \infty} ess sup \frac{d P_{n}}{d P} (X) = 1, \end{matrix}$

(31)

with $X \sim P$ . Then, (30) is locally tight in the sense that its both sides converge to 0, and

$\begin{matrix} lim_{n \to \infty} \frac{\frac{1}{3} χ^{2} (P_{n} ∥ P) + \frac{1}{6} χ^{2} (P ∥ P_{n})}{\frac{1}{log e} D (P_{n} ∥ P)} = 1 . \end{matrix}$

(32)
(c): For all $θ \in (0, 1)$ ,

$\begin{matrix} D (P ∥ Q) \geq (1 - θ) log (\frac{1}{1 - θ}) D_{ϕ_{θ}} (P ∥ Q) . \end{matrix}$

(33)

Moreover, under the assumption in (31), for all $θ \in [0, 1]$

$\begin{matrix} lim_{n \to \infty} \frac{D (P ∥ P_{n})}{D_{ϕ_{θ}} (P ∥ P_{n})} = \frac{1}{2} log e . \end{matrix}$

(34)
(d): [15] [Theorem 2]:

$\begin{matrix} \frac{1}{log e} D (P ∥ Q) \leq \frac{1}{2} χ^{2} (P ∥ Q) + \frac{1}{4} | P - Q | . \end{matrix}$

(35)

Proof.

See Section 5.2. □

Remark 2.

Inequality (30) is locally tight in the sense that (31) yields (32). This property, however, is not satisfied by (16) since the assumption in (31) implies that

\begin{matrix} lim_{n \to \infty} \frac{log (1 + χ^{2} (P_{n} ∥ P))}{D (P_{n} ∥ P)} = 2 . \end{matrix}

(36)

Remark 3.

Inequality (30) readily yields

\begin{matrix} D (P ∥ Q) + D (Q ∥ P) \leq \frac{1}{2} (χ^{2} (P ∥ Q) + χ^{2} (Q ∥ P)) log e, \end{matrix}

(37)

which is proved by a different approach in [27] [Proposition 4]. It is further shown in [15] [Theorem 2 b)] that

\begin{matrix} sup \frac{D (P ∥ Q) + D (Q ∥ P)}{χ^{2} (P ∥ Q) + χ^{2} (Q ∥ P)} = \frac{1}{2} log e, \end{matrix}

(38)

where the supremum is over

P ≪ ≫ Q

and

P \neq Q

.

3.2. Implications of Theorem 1

We next provide two implications of Theorem 1. The first implication, which relies on the Hammersley–Chapman–Robbins (HCR) bound for the chi-squared divergence [28,29], gives the following tight lower bound on the relative entropy

D (P ∥ Q)

as a function of the means and variances under P and Q.

Theorem 2.

Let P and Q be probability measures defined on the measurable space

(R, B)

, where

R

is the real line and

B

is the Borel σ–algebra of subsets of

R

. Let

m_{P}

,

m_{Q}

,

σ_{P}^{2}

, and

σ_{Q}^{2}

denote the expected values and variances of

X \sim P

and

Y \sim Q

, i.e.,

\begin{matrix} E [X] = : m_{P}, E [Y] = : m_{Q}, Var (X) = : σ_{P}^{2}, Var (Y) = : σ_{Q}^{2} . \end{matrix}

(39)

(a): If $m_{P} \neq m_{Q}$ , then

$\begin{matrix} D (P ∥ Q) \geq d (r ∥ s), \end{matrix}$

(40)

where $d (r ∥ s) : = r log \frac{r}{s} + (1 - r) log \frac{1 - r}{1 - s}$ , for $r, s \in [0, 1]$ , denotes the binary relative entropy (with the convention that $0 log \frac{0}{0} = 0$ ), and

$\begin{array}{l} (41) & r : = \frac{1}{2} + \frac{b}{4 a v} \in [0, 1], \\ (42) & s : = r - \frac{a}{2 v} \in [0, 1], \\ (43) & a : = m_{P} - m_{Q}, \\ (44) & b : = a^{2} + σ_{Q}^{2} - σ_{P}^{2}, \\ (45) & v : = \sqrt{σ_{P}^{2} + \frac{b^{2}}{4 a^{2}}} . \end{array}$
(b): The lower bound on the right side of (40) is attained for P and Q which are defined on the two-element set $U : = {u_{1}, u_{2}}$ , and

$\begin{matrix} P (u_{1}) = r, Q (u_{1}) = s, \end{matrix}$

(46)

with r and s in (41) and (42), respectively, and for $m_{P} \neq m_{Q}$

$\begin{matrix} u_{1} : = m_{P} + \sqrt{\frac{(1 - r) σ_{P}^{2}}{r}}, u_{2} : = m_{P} - \sqrt{\frac{r σ_{P}^{2}}{1 - r}} . \end{matrix}$

(47)
(c): If $m_{P} = m_{Q}$ and $σ_{P}$ and $σ_{Q}$ are selected arbitrarily, then

$\begin{matrix} inf_{P, Q} D (P ∥ Q) = 0, \end{matrix}$

(48)

where the infimum on the left side of (48) is taken over all P and Q which satisfy (39).

Proof.

See Section 5.3. □

Remark 4.

Consider the case of the non-equal means in Items (a) and (b) of Theorem 2. If these means are fixed, then the infimum of

D (P ∥ Q)

is zero by choosing arbitrarily large equal variances. Suppose now that the non-equal means

m_{P}

and

m_{Q}

are fixed, as well as one of the variances (either

σ_{P}^{2}

or

σ_{Q}^{2}

). Numerical experimentation shows that, in this case, the achievable lower bound in (40) is monotonically decreasing as a function of the other variance, and it tends to zero as we let the free variance tend to infinity. This asymptotic convergence to zero can be justified by assuming, for example, that

m_{P}, m_{Q}

, and

σ_{Q}^{2}

are fixed, and

m_{P} > m_{Q}

(the other cases can be justified in a similar way). Then, it can be verified from (41)–(45) that

\begin{matrix} r = \frac{{(m_{P} - m_{Q})}^{2}}{σ_{P}^{2}} + O (\frac{1}{σ_{P}^{4}}), s = O (\frac{1}{σ_{P}^{4}}), \end{matrix}

(49)

which implies that

d (r ∥ s) \to 0

as we let

σ_{P} \to \infty

. The infimum of the relative entropy

D (P ∥ Q)

is therefore equal to zero since the probability measures P and Q in (46) and (47), which are defined on a two-element set and attain the lower bound on the relative entropy under the constraints in (39), have a vanishing relative entropy in this asymptotic case.

Remark 5.

The proof of Item (c) in Theorem 2 suggests explicit constructions of sequences of pairs probability measures

{(P_{n}, Q_{n})}

such that

(a): The means under $P_{n}$ and $Q_{n}$ are both equal to m (independently of n);
(b): The variance under $P_{n}$ is equal to $σ_{P}^{2}$ , and the variance under $Q_{n}$ is equal to $σ_{Q}^{2}$ (independently of n);
(c): The relative entropy $D (P_{n} ∥ Q_{n})$ vanishes as we let $n \to \infty$ .

This yields in particular (48).

A second consequence of Theorem 1 gives the following result. Its first part holds due to the concavity of

exp (- D (P ∥ \cdot))

(see [30] [Problem 4.2]). The second part is new, and its proof relies on Theorem 1. As an educational note, we provide an alternative proof of the first part by relying on Theorem 1.

Theorem 3.

Let

P ≪ Q

, and

F : [0, 1] \to [0, \infty)

be given by

\begin{matrix} F (λ) : = D (P ∥ (1 - λ) P + λ Q), \forall λ \in [0, 1] . \end{matrix}

(50)

Then, for all

λ \in [0, 1]

,

\begin{matrix} F (λ) \leq log (\frac{1}{1 - λ + λ exp (- D (P ∥ Q))}), \end{matrix}

(51)

with an equality if

λ = 0

or

λ = 1

. Moreover, F is monotonically increasing, differentiable, and it satisfies

\begin{matrix} F^{'} (λ) \geq \frac{1}{λ} [exp (F (λ)) - 1] log e, \forall λ \in (0, 1], \end{matrix}

(52)

\begin{matrix} lim_{λ \to 0^{+}} \frac{F^{'} (λ)}{λ} = χ^{2} (Q ∥ P) log e, \end{matrix}

(53)

so the limit in (53) is twice as large as the value of the lower bound on this limit as it follows from the right side of (52).

Proof.

See Section 5.4. □

Remark 6.

By the convexity of the relative entropy, it follows that

F (λ) \leq λ D (P ∥ Q)

for all

λ \in [0, 1]

. It can be verified, however, that the inequality

1 - λ + λ exp (- x) \geq exp (- λ x)

holds for all

x \geq 0

and

λ \in [0, 1]

. Letting

x : = D (P ∥ Q)

implies that the upper bound on

F (λ)

on the right side of (51) is tighter than or equal to the upper bound

λ D (P ∥ Q)

(with an equality if and only if either

λ \in {0, 1}

or

P \equiv Q

).

Corollary 3.

Let

{P_{j}}_{j = 1}^{m}

, with

m \in N

, be probability measures defined on a measurable space

(X, F)

, and let

{α_{j}}_{j = 1}^{m}

be a sequence of non-negative numbers that sum to 1. Then, for all

i \in {1, \dots, m}

,

\begin{matrix} D (P_{i} ∥ \sum_{j = 1}^{m} α_{j} P_{j}) & \leq - log (α_{i} + (1 - α_{i}) exp (- \frac{1}{1 - α_{i}} \sum_{j \neq i} α_{j} D (P_{i} ∥ P_{j}))) . \end{matrix}

(54)

Proof.

For an arbitrary

i \in {1, \dots, m}

, apply the upper bound on the right side of (51) with

λ : = 1 - α_{i}

,

P : = P_{i}

and

Q : = \frac{1}{1 - α_{i}} \sum_{j \neq i} α_{j} P_{j}

. The right side of (54) is obtained from (51) by invoking the convexity of the relative entropy, which gives

D (P_{i} ∥ Q) \leq \frac{1}{1 - α_{i}} \sum_{j \neq i} α_{j} D (P_{i} ∥ P_{j})

. □

The next result provides an upper bound on the non-negative difference between the entropy of a convex combination of distributions and the respective convex combination of the individual entropies (it is also termed as the concavity deficit of the entropy function in [17] [Section 3]).

Corollary 4.

Let

{P_{j}}_{j = 1}^{m}

, with

m \in N

, be probability measures defined on a measurable space

(X, F)

, and let

{α_{j}}_{j = 1}^{m}

be a sequence of non-negative numbers that sum to 1. Then,

\begin{matrix} 0 \leq H (\sum_{j = 1}^{m} α_{j} P_{j}) - \sum_{j = 1}^{m} α_{j} H (P_{j}) \leq - \sum_{i = 1}^{m} α_{i} log (α_{i} + (1 - α_{i}) exp (- \frac{1}{1 - α_{i}} \sum_{j \neq i} α_{j} D (P_{i} ∥ P_{j}))) . \end{matrix}

(55)

Proof.

The lower bound holds due to the concavity of the entropy function. The upper bound readily follows from Corollary 3, and the identity

\begin{matrix} H (\sum_{j = 1}^{m} α_{j} P_{j}) - \sum_{j = 1}^{m} α_{j} H (P_{j}) = \sum_{i = 1}^{m} α_{i} D (P_{i} ∥ \sum_{j = 1}^{m} α_{j} P_{j}) . \end{matrix}

(56)

□

Remark 7.

The upper bound in (55) refines the known bound (see, e.g., [31] [Lemma 2.2])

\begin{matrix} H (\sum_{j = 1}^{m} α_{j} P_{j}) - \sum_{j = 1}^{m} α_{j} H (P_{j}) \leq \sum_{j = 1}^{m} α_{j} log \frac{1}{α_{j}} = H (\underset{̲}{α}), \end{matrix}

(57)

by relying on all the

\frac{1}{2} m (m - 1)

pairwise relative entropies between the individual distributions

{P_{j}}_{j = 1}^{m}

. Another refinement of (57), expressed in terms of total variation distances, has been recently provided in [17] [Theorem 3.1].

3.3. Monotonic Sequences of f-Divergences and an Extension of Theorem 1

The present subsection generalizes Theorem 1, and it also provides relations between f-divergences which are defined in a recursive way.

Theorem 4.

Let P and Q be probability measures defined on a measurable space

(X, F)

. Let

R_{λ}

, for

λ \in [0, 1]

, be the convex combination of P and Q as in (21). Let

f_{0} : (0, \infty) \to R

be a convex function with

f_{0} (1) = 0

, and let

{f_{k} (\cdot)}_{k = 0}^{\infty}

be a sequence of functions that are defined on

(0, \infty)

by the recursive equation

\begin{matrix} f_{k + 1} (x) : = \int_{0}^{1 - x} f_{k} (1 - s) \frac{d s}{s}, x > 0, k \in {0, 1, \dots} . \end{matrix}

(58)

Then,

(a): ${\{D_{f_{k}} (P ∥ Q)\}}_{k = 0}^{\infty}$ is a non-increasing (and non-negative) sequence of f-divergences.
(b): For all $λ \in [0, 1]$ and $k \in {0, 1, \dots}$ ,

$\begin{matrix} D_{f_{k + 1}} (R_{λ} ∥ P) = \int_{0}^{λ} D_{f_{k}} (R_{s} ∥ P) \frac{d s}{s} . \end{matrix}$

(59)

Proof.

See Section 5.5. □

We next use the polylogarithm functions, which satisfy the recursive equation [32] [Equation (7.2)]:

{Li}_{k} (x) : = {\begin{cases} \frac{x}{1 - x}, & i f k = 0, \\ \int_{0}^{x} \frac{{Li}_{k - 1} (s)}{s} d s, & i f k \geq 1 . \end{cases}

(60)

This gives

{Li}_{1} (x) = - {log}_{e} (1 - x)

,

{Li}_{2} (x) = - \int_{0}^{x} \frac{1}{s} {log}_{e} (1 - s) d s

and so on, which are real–valued and finite for

x < 1

.

Corollary 5.

Let

\begin{matrix} f_{k} (x) : = {Li}_{k} (1 - x), x > 0, k \in {0, 1, \dots} . \end{matrix}

(61)

Then, (59) holds for all

λ \in [0, 1]

and

k \in {0, 1, \dots}

. Furthermore, setting

k = 0

in (59) yields (22) as a special case.

Proof.

See Section 5.6. □

3.4. On Probabilities and f-Divergences

The following result relates probabilities of sets to f-divergences.

Theorem 5.

Let

(X, F, μ)

be a probability space, and let

C \in F

be a measurable set with

μ (C) > 0

. Define the conditional probability measure

\begin{matrix} μ_{C} (E) : = \frac{μ (C \cap E)}{μ (C)}, \forall E \in F . \end{matrix}

(62)

Let

f : (0, \infty) \to R

be an arbitrary convex function with

f (1) = 0

, and assume (by continuous extension of f at zero) that

f (0) : = lim_{t \to 0^{+}} f (t) < \infty

. Furthermore, let

\tilde{f} : (0, \infty) \to R

be the convex function which is given by

\begin{matrix} \tilde{f} (t) : = t f (\frac{1}{t}), \forall t > 0 . \end{matrix}

(63)

Then,

\begin{matrix} D_{f} (μ_{C} ∥ μ) = \tilde{f} (μ (C)) + (1 - μ (C)) f (0) . \end{matrix}

(64)

Proof.

See Section 5.7. □

Connections of probabilities to the relative entropy, and to the chi-squared divergence, are next exemplified as special cases of Theorem 5.

Corollary 6.

In the setting of Theorem 5,

\begin{matrix} D (μ_{C} ∥ μ) = log \frac{1}{μ (C)}, \end{matrix}

(65)

\begin{matrix} χ^{2} (μ_{C} ∥ μ) = \frac{1}{μ (C)} - 1, \end{matrix}

(66)

so (16) is satisfied in this case with equality. More generally, for all

α \in (0, \infty)

,

\begin{matrix} D_{α} (μ_{C} ∥ μ) = log \frac{1}{μ (C)} . \end{matrix}

(67)

Proof.

See Section 5.7. □

Remark 8.

In spite of its simplicity, (65) proved very useful in the seminal work by Marton on transportation–cost inequalities, proving concentration of measures by information-theoretic tools [33,34] (see also [35] [Chapter 8] and [36] [Chapter 3]). As a side note, the simple identity (65) was apparently first explicitly used by Csiszár (see [37] [Equation (4.13)]).

4. Applications

This section provides applications of our results in Section 3. These include universal lossless compression, method of types and large deviations, and strong data–processing inequalities (SDPIs).

4.1. Application of Corollary 3: Shannon Code for Universal Lossless Compression

Consider

m > 1

discrete, memoryless, and stationary sources with probability mass functions

{P_{i}}_{i = 1}^{m}

, and assume that the symbols are emitted by one of these sources with an a priori probability

α_{i}

for source no. i, where

{α_{i}}_{i = 1}^{m}

are positive and sum to 1.

For lossless data compression by a universal source code, suppose that a single source code is designed with respect to the average probability mass function

P : = \sum_{j = 1}^{m} α_{j} P_{j}

.

Assume that the designer uses a Shannon code, where the code assignment for a symbol

x \in X

is of length

ℓ (x) = ⌈ log \frac{1}{P (x)} ⌉

bits (logarithms are on base 2). Due to the mismatch in the source distribution, the average codeword length

ℓ_{avg}

satisfies (see [38] [Proposition 3.B])

\begin{matrix} \sum_{i = 1}^{m} α_{i} H (P_{i}) + \sum_{i = 1}^{m} α_{i} D (P_{i} ∥ P) \leq ℓ_{avg} \leq \sum_{i = 1}^{m} α_{i} H (P_{i}) + \sum_{i = 1}^{m} α_{i} D (P_{i} ∥ P) + 1 . \end{matrix}

(68)

The fractional penalty in the average codeword length, denoted by

ν

, is defined to be equal to the ratio of the penalty in the average codeword length as a result of the source mismatch, and the average codeword length in case of a perfect matching. From (68), it follows that

\begin{matrix} \frac{\sum_{j = 1}^{m} α_{i} D (P_{i} ∥ P)}{1 + \sum_{j = 1}^{m} α_{i} H (P_{i})} \leq ν \leq \frac{1 + \sum_{j = 1}^{m} α_{i} D (P_{i} ∥ P)}{\sum_{j = 1}^{m} α_{i} H (P_{i})} . \end{matrix}

(69)

We next rely on Corollary 3 to obtain an upper bound on

ν

which is expressed as a function of the

m (m - 1)

relative entropies

D (P_{i} ∥ P_{j})

for all

i \neq j

in

{1, \dots, m}

. This is useful if, e.g., the m relative entropies on the left and right sides of (69) do not admit closed form expressions, in contrast to the

m (m - 1)

relative entropies

D (P_{i} ∥ P_{j})

for

i \neq j

. We next exemplify this case.

For

i \in {1, \dots, m}

, let

P_{i}

be a Poisson distribution with parameter

λ_{i} > 0

. For all

i, j \in {1, \dots, m}

, the relative entropy from

P_{i}

to

P_{j}

admits the closed-form expression

\begin{matrix} D (P_{i} ∥ P_{j}) = λ_{i} log (\frac{λ_{i}}{λ_{j}}) + (λ_{j} - λ_{i}) log e . \end{matrix}

(70)

From (54) and (70), it follows that

\begin{matrix} D (P_{i} ∥ P) & \leq - log (α_{i} + (1 - α_{i}) exp (- \frac{f_{i} (\underset{̲}{α}, \underset{̲}{λ})}{1 - α_{i}})), \end{matrix}

(71)

where

\begin{matrix} (72) & f_{i} (\underset{̲}{α}, \underset{̲}{λ}) & : = \sum_{j \neq i} α_{j} D (P_{i} ∥ P_{j}) \\ (73) & = \sum_{j \neq i} \{α_{j} [λ_{i} log (\frac{λ_{i}}{λ_{j}}) + (λ_{j} - λ_{i}) log e]\} . \end{matrix}

The entropy of a Poisson distribution, with parameter

λ_{i}

, is given by the integral representation [39,40,41]

\begin{matrix} H (P_{i}) = λ_{i} log (\frac{e}{λ_{i}}) + \int_{0}^{\infty} (λ_{i} - \frac{1 - e^{- λ_{i} (1 - e^{- u})}}{1 - e^{- u}}) \frac{e^{- u}}{u} d u log e . \end{matrix}

(74)

Combining (69), (71) and (74) finally gives an upper bound on

ν

in the considered setup.

Example 1.

Consider five discrete memoryless sources where the probability mass function of source no. i is given by

P_{i} = Poisson (λ_{i})

with

\underset{̲}{λ} = [16, 20, 24, 28, 32]

. Suppose that the symbols are emitted from one of the sources with equal probability, so

\underset{̲}{α} = [\frac{1}{5}, \frac{1}{5}, \frac{1}{5}, \frac{1}{5}, \frac{1}{5}]

. Let

P : = \frac{1}{5} (P_{1} + \dots + P_{5})

be the average probability mass function of the five sources. The term

\sum_{i} α_{i} D (P_{i} ∥ P)

, which appears in the numerators of the upper and lower bounds on ν (see (69)), does not lend itself to a closed-form expression, and it is not even an easy task to calculate it numerically due to the need to compute an infinite series which involves factorials. We therefore apply the closed-form upper bound in (71) to get that

\sum_{i} α_{i} D (P_{i} ∥ P) \leq 1.46

bits, whereas the upper bound which follows from the convexity of the relative entropy (i.e.,

\sum_{i} α_{i} f_{i} (\underset{̲}{α}, \underset{̲}{λ})

) is equal to 1.99 bits (both upper bounds are smaller than the trivial bound

{log}_{2} 5 \approx 2.32

bits). From (69), (74), and the stronger upper bound on

\sum_{i} α_{i} D (P_{i} ∥ P)

, the improved upper bound on ν is equal to

57.0 %

(as compared to a looser upper bound of

69.3 %

, which follows from (69), (74), and the looser upper bound on

\sum_{i} α_{i} D (P_{i} ∥ P)

that is equal to 1.99 bits).

4.2. Application of Theorem 2 in the Context of the Method of Types and Large Deviations Theory

Let

X^{n} = (X_{1}, \dots, X_{n})

be a sequence of i.i.d. random variables with

X_{1} \sim Q

, where Q is a probability measure defined on a finite set

X

, and

Q (x) > 0

for all

x \in X

. Let

P

be a set of probability measures on

X

such that

Q \notin P

, and suppose that the closure of

P

coincides with the closure of its interior. Then, by Sanov’s theorem (see, e.g., [42] [Theorem 11.4.1] and [43] [Theorem 3.3]), the probability that the empirical distribution

{\hat{P}}_{X^{n}}

belongs to

P

vanishes exponentially at the rate

\begin{matrix} lim_{n \to \infty} \frac{1}{n} log \frac{1}{P [{\hat{P}}_{X^{n}} \in P]} = inf_{P \in P} D (P ∥ Q) . \end{matrix}

(75)

Furthermore, for finite n, the method of types yields the following upper bound on this rare event:

\begin{matrix} (76) & P [{\hat{P}}_{X^{n}} \in P] & \leq (\binom{n + | X | - 1}{| X | - 1}) exp (- n inf_{P \in P} D (P ∥ Q)) \\ (77) & \leq {(n + 1)}^{| X | - 1} exp (- n inf_{P \in P} D (P ∥ Q)), \end{matrix}

whose exponential decay rate coincides with the exact asymptotic result in (75).

Suppose that Q is not fully known, but its mean

m_{Q}

and variance

σ_{Q}^{2}

are available. Let

m_{1} \in R

and

δ_{1}, ε_{1}, σ_{1} > 0

be fixed, and let

P

be the set of all probability measures P, defined on the finite set

X

, with mean

m_{P} \in [m_{1} - δ_{1}, m_{1} + δ_{1}]

and variance

σ_{P}^{2} \in [σ_{1}^{2} - ε_{1}, σ_{1}^{2} + ε_{1}]

, where

| m_{1} - m_{Q} | > δ_{1}

. Hence,

P

coincides with the closure of its interior, and

Q \notin P

.

The lower bound on the relative entropy in Theorem 2, used in conjunction with the upper bound in (77), can serve to obtain an upper bound on the probability of the event that the empirical distribution of

X^{n}

belongs to the set

P

, regardless of the uncertainty in Q. This gives

\begin{matrix} P [{\hat{P}}_{X^{n}} \in P] & \leq {(n + 1)}^{| X | - 1} exp (- n d^{*}), \end{matrix}

(78)

where

\begin{matrix} d^{*} : = inf_{m_{P}, σ_{P}^{2}} d (r ∥ s), \end{matrix}

(79)

and, for fixed

(m_{P}, m_{Q}, σ_{P}^{2}, σ_{Q}^{2})

, the parameters r and s are given in (41) and (42), respectively.

Standard algebraic manipulations that rely on (78) lead to the following result, which is expressed as a function of the Lambert–W function [44]. This function, which finds applications in various engineering and scientific fields, is a standard built–in function in mathematical software tools such as Mathematica, Matlab, and Maple. Applications of the Lambert–W function in information theory and coding are briefly surveyed in [45].

Proposition 2.

For

ε \in (0, 1)

, let

n^{*} : = n^{*} (ε)

denote the minimal value of

n \in N

such that the upper bound on the right side of (78) does not exceed

ε \in (0, 1)

. Then,

n^{*}

admits the following closed-form expression:

\begin{matrix} n^{*} = max \{⌈- \frac{(| X | - 1) W_{- 1} (η) log e}{d^{*}}⌉ - 1, 1\}, \end{matrix}

(80)

with

\begin{matrix} η : = - \frac{d^{*} {(ε exp (- d^{*}))}^{1 / (| X | - 1)}}{(| X | - 1) log e} \in [- \frac{1}{e}, 0), \end{matrix}

(81)

and

W_{- 1} (\cdot)

on the right side of (80) denotes the secondary real–valued branch of the Lambert–W function (i.e.,

x : = W_{- 1} (y)

where

W_{- 1} : [- \frac{1}{e}, 0) \to (- \infty, - 1]

is the inverse function of

y : = x e^{x}

).

Example 2.

Let Q be an arbitrary probability measure, defined on a finite set

X

, with mean

m_{Q} = 40

and variance

σ_{Q}^{2} = 20

. Let

P

be the set of all probability measures P, defined on

X

, whose mean

m_{P}

and variance

σ_{P}^{2}

lie in the intervals

[43, 47]

and

[18, 22]

, respectively. Suppose that it is required that, for all probability measures Q as above, the probability that the empirical distribution of the i.i.d. sequence

X^{n} \sim Q^{n}

that is included in the set

P

is at most

ε = 10^{- 10}

. We rely here on the upper bound in (78), and impose the stronger condition where it should not exceed ε. By this approach, it is obtained numerically from (79) that

d^{*} = 0.203

nats. We next examine two cases:

(i): If $| X | = 2$ , then it follows from (80) that $n^{*} = 138$ .
(ii): Consider a richer alphabet size of the i.i.d. samples where, e.g., $| X | = 100$ . By relying on the same universal lower bound $d^{*}$ , which holds independently of the value of $| X |$ ( $X$ can possibly be an infinite set), it follows from (80) that $n^{*} = 4170$ is the minimal value such that the upper bound in (78) does not exceed $10^{- 10}$ .

We close this discussion by providing numerical experimentation of the lower bound on the relative entropy in Theorem 2, and comparing this attainable lower bound (see Item b) of Theorem 2) with the following closed-form expressions for relative entropies:

(a): The relative entropy between real-valued Gaussian distributions is given by

$\begin{matrix} D (N (m_{P}, σ_{P}^{2}) ∥ N (m_{Q}, σ_{Q}^{2})) = log \frac{σ_{Q}}{σ_{P}} + \frac{1}{2} [\frac{{(m_{P} - m_{Q})}^{2} + σ_{P}^{2}}{σ_{Q}^{2}} - 1] log e . \end{matrix}$

(82)
(b): Let $E_{μ}$ denote a random variable which is exponentially distributed with mean $μ > 0$ ; its probability density function is given by

$\begin{matrix} e_{μ} (x) = \frac{1}{μ} e^{- x / μ} 1 {x \geq 0} . \end{matrix}$

(83)

Then, for $a_{1}, a_{2} > 0$ and $d_{1}, d_{2} \in R$ ,

$D (E_{a_{1}} + d_{1} ∥ E_{a_{2}} + d_{2}) = {\begin{cases} log \frac{a_{2}}{a_{1}} + \frac{d_{1} + a_{1} - d_{2} - a_{2}}{a_{2}} log e, & d_{1} \geq d_{2}, \\ \infty, & d_{1} < d_{2} . \end{cases}$

(84)

In this case, the means under P and Q are $m_{P} = d_{1} + a_{1}$ and $m_{Q} = d_{2} + a_{2}$ , respectively, and the variances are $σ_{P}^{2} = a_{1}^{2}$ and $σ_{Q}^{2} = a_{2}^{2}$ . Hence, for obtaining the required means and variances, set

$\begin{matrix} a_{1} = σ_{P}, a_{2} = σ_{Q}, d_{1} = m_{P} - σ_{P}, d_{2} = m_{Q} - σ_{Q} . \end{matrix}$

(85)

Example 3.

We compare numerically the attainable lower bound on the relative entropy, as it given in (40), with the two relative entropies in (82) and (84):

(i): If $(m_{P}, m_{Q}, σ_{P}^{2}, σ_{Q}^{2}) = (45, 40, 20, 20)$ , then the lower bound in (40) is equal to 0.521 nats, and the two relative entropies in (82) and (84) are equal to 0.625 and 1.118 nats, respectively.
(ii): If $(m_{P}, m_{Q}, σ_{P}^{2}, σ_{Q}^{2}) = (50, 35, 10, 20)$ , then the lower bound in (40) is equal to 2.332 nats, and the two relative entropies in (82) and (84) are equal to 5.722 and 3.701 nats, respectively.

4.3. Strong Data–Processing Inequalities and Maximal Correlation

The information contraction is a fundamental concept in information theory. The contraction of f-divergences through channels is captured by data–processing inequalities, which can be further tightened by the derivation of SDPIs with channel-dependent or source-channel dependent contraction coefficients (see, e.g., [26,46,47,48,49,50,51,52]).

We next provide necessary definitions which are relevant for the presentation in this subsection.

Definition 7.

Let

Q_{X}

be a probability distribution which is defined on a set

X

, and that is not a point mass, and let

W_{Y | X} : X \to Y

be a stochastic transformation. The contraction coefficient for f-divergences is defined as

\begin{matrix} μ_{f} (Q_{X}, W_{Y | X}) : = sup_{P_{X} : D_{f} (P_{X} ∥ Q_{X}) \in (0, \infty)} \frac{D_{f} (P_{Y} ∥ Q_{Y})}{D_{f} (P_{X} ∥ Q_{X})}, \end{matrix}

(86)

where, for all

y \in Y

,

\begin{matrix} P_{Y} (y) = (P_{X} W_{Y | X}) (y) : = \int_{X} d P_{X} (x) W_{Y | X} (y | x), \end{matrix}

(87)

\begin{matrix} Q_{Y} (y) = (Q_{X} W_{Y | X}) (y) : = \int_{X} d Q_{X} (x) W_{Y | X} (y | x) . \end{matrix}

(88)

The notation in (87) and (88) is consistent with the standard notation used in information theory (see, e.g., the first displayed equation after (3.2) in [53]).

The derivation of good upper bounds on contraction coefficients for f-divergences, which are strictly smaller than 1, lead to SDPIs. These inequalities find their applications, e.g., in studying the exponential convergence rate of an irreducible, time-homogeneous and reversible discrete-time Markov chain to its unique invariant distribution over its state space (see, e.g., [49] [Section 2.4.3] and [50] [Section 2]). It is in sharp contrast to DPIs which do not yield convergence to stationarity at any rate. We return to this point later in this subsection, and determine the exact convergence rate to stationarity under two parametric families of f-divergences.

We next rely on Theorem 1 to obtain upper bounds on the contraction coefficients for the following f-divergences.

Definition 8.

For

α \in (0, 1]

, the α-skew K-divergence is given by

\begin{matrix} K_{α} (P ∥ Q) : = D (P ∥ (1 - α) P + α Q), \end{matrix}

(89)

and, for

α \in [0, 1]

, let

\begin{matrix} (90) & S_{α} (P ∥ Q) & : = α D (P ∥ (1 - α) P + α Q) + (1 - α) D (Q ∥ (1 - α) P + α Q) \\ (91) & = α K_{α} (P ∥ Q) + (1 - α) K_{1 - α} (Q ∥ P), \end{matrix}

with the convention that

K_{0} (P ∥ Q) \equiv 0

(by a continuous extension at

α = 0

in (89)). These divergence measures are specialized to the relative entropies:

\begin{matrix} K_{1} (P ∥ Q) = D (P ∥ Q) = S_{1} (P ∥ Q), & S_{0} (P ∥ Q) = D (Q ∥ P), \end{matrix}

(92)

and

S_{\frac{1}{2}} (P ∥ Q)

is the Jensen–Shannon divergence [54,55,56] (also known as the capacitory discrimination [57]):

\begin{matrix} (93) & S_{\frac{1}{2}} (P ∥ Q) & = \frac{1}{2} D (P ∥ \frac{1}{2} (P + Q)) + \frac{1}{2} D (Q ∥ \frac{1}{2} (P + Q)) \\ (94) & = H (\frac{1}{2} (P + Q)) - \frac{1}{2} H (P) - \frac{1}{2} H (Q) : = JS (P ∥ Q) . \end{matrix}

It can be verified that the divergence measures in (89) and (90) are f-divergences:

\begin{matrix} K_{α} (P ∥ Q) = D_{k_{α}} (P ∥ Q), α \in (0, 1], \end{matrix}

(95)

\begin{matrix} S_{α} (P ∥ Q) = D_{s_{α}} (P ∥ Q), α \in [0, 1], \end{matrix}

(96)

with

\begin{matrix} (97) & k_{α} (t) & : = t log t - t log (α + (1 - α) t), t > 0, α \in (0, 1], \\ (98) & s_{α} (t) & : = α t log t - (α t + 1 - α) log (α + (1 - α) t) \\ (99) & = α k_{α} (t) + (1 - α) t k_{1 - α} (\frac{1}{t}), t > 0, α \in [0, 1], \end{matrix}

where

k_{α} (\cdot)

and

s_{α} (\cdot)

are strictly convex functions on

(0, \infty)

, and vanish at 1.

Remark 9.

The α-skew K-divergence in (89) is considered in [55] and [58] [(13)] (including pointers in the latter paper to its utility). The divergence in (90) is akin to Lin’s measure in [55] [(4.1)], the asymmetric α-skew Jensen–Shannon divergence in [58] [(11)–(12)], the symmetric α-skew Jensen–Shannon divergence in [58] [(16)], and divergence measures in [59] which involve arithmetic and geometric means of two probability distributions. Properties and applications of quantum skew divergences are studied in [19] and references therein.

Theorem 6.

The f-divergences in (89) and (90) satisfy the following integral identities, which are expressed in terms of the Györfi–Vajda divergence in (17):

\begin{matrix} \frac{1}{log e} K_{α} (P ∥ Q) = \int_{0}^{α} s D_{ϕ_{s}} (P ∥ Q) d s, α \in (0, 1], \end{matrix}

(100)

\begin{matrix} \frac{1}{log e} S_{α} (P ∥ Q) = \int_{0}^{1} g_{α} (s) D_{ϕ_{s}} (P ∥ Q) d s, α \in [0, 1], \end{matrix}

(101)

with

\begin{matrix} g_{α} (s) : = α s 1 \{s \in (0, α]\} + (1 - α) (1 - s) 1 \{s \in [α, 1)\}, (α, s) \in [0, 1] \times [0, 1] . \end{matrix}

(102)

Moreover, the contraction coefficients for these f-divergences are related as follows:

\begin{matrix} μ_{χ^{2}} (Q_{X}, W_{Y | X}) & \leq μ_{k_{α}} (Q_{X}, W_{Y | X}) \leq sup_{s \in (0, α]} μ_{ϕ_{s}} (Q_{X}, W_{Y | X}), α \in (0, 1], \end{matrix}

(103)

\begin{matrix} μ_{χ^{2}} (Q_{X}, W_{Y | X}) & \leq μ_{s_{α}} (Q_{X}, W_{Y | X}) \leq sup_{s \in (0, 1)} μ_{ϕ_{s}} (Q_{X}, W_{Y | X}), α \in [0, 1], \end{matrix}

(104)

where

μ_{χ^{2}} (Q_{X}, W_{Y | X})

denotes the contraction coefficient for the chi-squared divergence.

Proof.

See Section 5.8. □

Remark 10.

The upper bounds on the contraction coefficients for the parametric f-divergences in (89) and (90) generalize the upper bound on the contraction coefficient for the relative entropy in [51] [Theorem III.6] (recall that

K_{1} (P ∥ Q) = D (P ∥ Q) = S_{1} (P ∥ Q)

), so the upper bounds in Theorem 6 are specialized to the latter bound at

α = 1

.

Corollary 7.

Let

\begin{matrix} μ_{χ^{2}} (W_{Y | X}) : = sup_{Q} μ_{χ^{2}} (Q_{X}, W_{Y | X}), \end{matrix}

(105)

where the supremum on the right side is over all probability measures

Q_{X}

defined on

X

. Then,

\begin{matrix} μ_{χ^{2}} (Q_{X}, W_{Y | X}) & \leq μ_{k_{α}} (Q_{X}, W_{Y | X}) \leq μ_{χ^{2}} (W_{Y | X}), α \in (0, 1], \end{matrix}

(106)

\begin{matrix} μ_{χ^{2}} (Q_{X}, W_{Y | X}) & \leq μ_{s_{α}} (Q_{X}, W_{Y | X}) \leq μ_{χ^{2}} (W_{Y | X}), α \in [0, 1] . \end{matrix}

(107)

Proof.

See Section 5.9. □

Example 4.

Let

Q_{X} = Bernoulli (\frac{1}{2})

, and let

W_{Y | X}

correspond to a binary symmetric channel (BSC) with crossover probability ε. Then,

μ_{χ^{2}} (Q_{X}, W_{Y | X}) = μ_{χ^{2}} (W_{Y | X}) = {(1 - 2 ε)}^{2}

. The upper and lower bounds on

μ_{k_{α}} (Q_{X}, W_{Y | X})

and

μ_{s_{α}} (Q_{X}, W_{Y | X})

in (106) and (107) match for all α, and they are all equal to

{(1 - 2 ε)}^{2}

.

The upper bound on the contraction coefficients in Corollary 7 is given by

μ_{χ^{2}} (W_{Y | X})

, whereas the lower bound is given by

μ_{χ^{2}} (Q_{X}, W_{Y | X})

, which depends on the input distribution

Q_{X}

. We next provide alternative upper bounds on the contraction coefficients for the considered (parametric) f-divergences, which, similarly to the lower bound, scale like

μ_{χ^{2}} (Q_{X}, W_{Y | X})

. Although the upper bound in Corollary 7 may be tighter in some cases than the alternative upper bounds which are next presented in Proposition 3 (and in fact, the former upper bound may be even achieved with equality as in Example 4), the bounds in Proposition 3 are used shortly to determine the exponential rate of the convergence to stationarity of a type of Markov chains.

Proposition 3.

For all

α \in (0, 1]

,

\begin{matrix} (108) & μ_{χ^{2}} (Q_{X}, W_{Y | X}) \leq μ_{k_{α}} (Q_{X}, W_{Y | X}) \leq \frac{1}{α Q_{min}} \cdot μ_{χ^{2}} (Q_{X}, W_{Y | X}), \\ (109) & μ_{χ^{2}} (Q_{X}, W_{Y | X}) \leq μ_{s_{α}} (Q_{X}, W_{Y | X}) \leq \frac{(1 - α) {log}_{e} (\frac{1}{α}) + 2 α - 1}{(1 - 3 α + 3 α^{2}) Q_{min}} \cdot μ_{χ^{2}} (Q_{X}, W_{Y | X}), \end{matrix}

where

Q_{min}

denotes the minimal positive mass of the input distribution

Q_{X}

.

Proof.

See Section 5.10. □

Remark 11.

In view of (92), at

α = 1

, (108) and (109) specialize to an upper bound on the contraction coefficient of the relative entropy (KL divergence) as a function of the contraction coefficient of the chi-squared divergence. In this special case, both (108) and (109) give

\begin{matrix} μ_{χ^{2}} (Q_{X}, W_{Y | X}) \leq μ_{KL} (Q_{X}, W_{Y | X}) \leq \frac{1}{Q_{min}} \cdot μ_{χ^{2}} (Q_{X}, W_{Y | X}), \end{matrix}

(110)

which then coincides with [48] [Theorem 10].

We next apply Proposition 3 to consider the convergence rate to stationarity of Markov chains by the introduced f-divergences in Definition 8. The next result follows [49] [Section 2.4.3], and it provides a generalization of the result there.

Theorem 7.

Consider a time-homogeneous, irreducible, and reversible discrete-time Markov chain with a finite state space

X

, let W be its probability transition matrix, and

Q_{X}

be its unique stationary distribution (reversibility means that

Q_{X} (x) {[W]}_{x, y} = Q_{X} (y) {[W]}_{y, x}

for all

x, y \in X

). Let

P_{X}

be an initial probability distribution over

X

. Then, for all

α \in (0, 1]

and

n \in N

,

\begin{matrix} (111) & K_{α} (P_{X} W^{n} ∥ Q_{X}) \leq μ_{k_{α}} (Q_{X}, W^{n}) K_{α} (P_{X} ∥ Q_{X}), \\ (112) & S_{α} (P_{X} W^{n} ∥ Q_{X}) \leq μ_{s_{α}} (Q_{X}, W^{n}) S_{α} (P_{X} ∥ Q_{X}), \end{matrix}

and the contraction coefficients on the right sides of (111) and (112) scale like the n-th power of the contraction coefficient for the chi-squared divergence as follows:

\begin{matrix} (113) & {(μ_{χ^{2}} (Q_{X}, W))}^{n} \leq μ_{k_{α}} (Q_{X}, W^{n}) \leq \frac{1}{α Q_{min}} \cdot {(μ_{χ^{2}} (Q_{X}, W))}^{n}, \\ (114) & {(μ_{χ^{2}} (Q_{X}, W))}^{n} \leq μ_{s_{α}} (Q_{X}, W^{n}) \leq \frac{(1 - α) {log}_{e} (\frac{1}{α}) + 2 α - 1}{(1 - 3 α + 3 α^{2}) Q_{min}} \cdot {(μ_{χ^{2}} (Q_{X}, W))}^{n} . \end{matrix}

Proof.

Inequalities (111) and (112) hold since

Q_{X} W^{n} = Q_{X}

, for all

n \in N

, and due to Definition 7 and (95) and (96). Inequalities (113) and (114) hold by Proposition 3, and due to the reversibility of the Markov chain which implies that (see [49] [Equation (2.92)])

\begin{matrix} μ_{χ^{2}} (Q_{X}, W^{n}) = {(μ_{χ^{2}} (Q_{X}, W))}^{n}, n \in N . \end{matrix}

(115)

□

In view of (113) and (114), Theorem 7 readily gives the following result on the exponential decay rate of the upper bounds on the divergences on the left sides of (111) and (112).

Corollary 8.

For all

α \in (0, 1]

,

\begin{matrix} lim_{n \to \infty} {(μ_{k_{α}} (Q_{X}, W^{n}))}^{1 / n} = μ_{χ^{2}} (Q_{X}, W) = lim_{n \to \infty} {(μ_{s_{α}} (Q_{X}, W^{n}))}^{1 / n} . \end{matrix}

(116)

Remark 12.

Theorem 7 and Corollary 8 generalize the results in [49] [Section 2.4.3], which follow as a special case at

α = 1

(see (92)).

We end this subsection by considering maximal correlations, which are closely related to the contraction coefficient for the chi-squared divergence.

Definition 9.

The maximal correlation between two random variables X and Y is defined as

\begin{matrix} ρ_{m} (X; Y) : = sup_{f, g} E [f (X) g (Y)], \end{matrix}

(117)

where the supremum is taken over all real-valued functions f and g such that

\begin{matrix} E [f (X)] = E [g (Y)] = 0, E [f^{2} (X)] \leq 1, E [g^{2} (Y)] \leq 1 . \end{matrix}

(118)

It is well-known [60] that, if

X \sim Q_{X}

and

Y \sim Q_{Y} = Q_{X} W_{Y | X}

, then the contraction coefficient for the chi-squared divergence

μ_{χ^{2}} (Q_{X}, W_{Y | X})

is equal to the square of the maximal correlation between the random variables X and Y, i.e.,

\begin{matrix} ρ_{m} (X; Y) = \sqrt{μ_{χ^{2}} (Q_{X}, W_{Y | X})} . \end{matrix}

(119)

A simple application of Corollary 1 and (119) gives the following result.

Proposition 4.

In the setting of Definition 7, for

s \in [0, 1]

, let

X_{s} \sim (1 - s) P_{X} + s Q_{X}

and

Y_{s} \sim (1 - s) P_{Y} + s Q_{Y}

with

P_{X} \neq Q_{X}

and

P_{X} ≪ ≫ Q_{X}

. Then, the following inequality holds:

\begin{matrix} sup_{s \in [0, 1]} ρ_{m} (X_{s}; Y_{s}) \geq max \{\sqrt{\frac{D (P_{Y} ∥ Q_{Y})}{D (P_{X} ∥ Q_{X})}}, \sqrt{\frac{D (Q_{Y} ∥ P_{Y})}{D (Q_{X} ∥ P_{X})}}\} . \end{matrix}

(120)

Proof.

See Section 5.11. □

5. Proofs

This section provides proofs of the results in Section 3 and Section 4.

5.1. Proof of Theorem 1

Proof of (22): We rely on an integral representation of the logarithm function (on base

e

):

\begin{matrix} {log}_{e} x & = \int_{0}^{1} \frac{x - 1}{x + (1 - x) v} d v, \forall x > 0 . \end{matrix}

(121)

Let

μ

be a dominating measure of P and Q (i.e.,

P, Q ≪ μ

), and let

p : = \frac{d P}{d μ}

,

q : = \frac{d Q}{d μ}

, and

\begin{matrix} r_{λ} : = \frac{d R_{λ}}{d μ} = (1 - λ) p + λ q, \forall λ \in [0, 1], \end{matrix}

(122)

where the last equality is due to (21). For all

λ \in [0, 1]

,

\begin{matrix} (123) & \frac{1}{log e} D (P ∥ R_{λ}) & = \int p {log}_{e} (\frac{p}{r_{λ}}) d μ \\ (124) & = \int_{0}^{1} \int \frac{p (p - r_{λ})}{p + v (r_{λ} - p)} d μ d v, \end{matrix}

where (124) holds due to (121) with

x : = \frac{p}{r_{λ}}

, and by swapping the order of integration. The inner integral on the right side of (124) satisfies, for all

v \in (0, 1]

,

\begin{matrix} (125) & \int \frac{p (p - r_{λ})}{p + v (r_{λ} - p)} d μ & = \int (p - r_{λ}) (1 + \frac{v (p - r_{λ})}{p + v (r_{λ} - p)}) d μ \\ (126) & = \int (p - r_{λ}) d μ + v \int \frac{{(p - r_{λ})}^{2}}{p + v (r_{λ} - p)} d μ \\ (127) & = v \int \frac{{(p - r_{λ})}^{2}}{(1 - v) p + v r_{λ}} d μ \\ (128) & = \frac{1}{v} \int \frac{{(p - [(1 - v) p + v r_{λ}])}^{2}}{(1 - v) p + v r_{λ}} d μ \\ (129) & = \frac{1}{v} χ^{2} (P ∥ (1 - v) P + v R_{λ}), \end{matrix}

where (127) holds since

\int p d μ = 1

, and

\int r_{λ} d μ = 1

. From (21), for all

(λ, v) \in [0, 1] \times [0, 1]

,

\begin{matrix} (1 - v) P + v R_{λ} = (1 - λ v) P + λ v Q = R_{λ v} . \end{matrix}

(130)

The substitution of (130) into the right side of (129) gives that, for all

(λ, v) \in [0, 1] \times (0, 1]

,

\begin{matrix} \int \frac{p (p - r_{λ})}{p + v (r_{λ} - p)} d μ = \frac{1}{v} χ^{2} (P ∥ R_{λ v}) . \end{matrix}

(131)

Finally, substituting (131) into the right side of (124) gives that, for all

λ \in (0, 1]

,

\begin{matrix} (132) & \frac{1}{log e} D (P ∥ R_{λ}) & = \int_{0}^{1} \frac{1}{v} χ^{2} (P ∥ R_{λ v}) d v \\ (133) & = \int_{0}^{λ} \frac{1}{s} χ^{2} (P ∥ R_{s}) d s, \end{matrix}

where (133) holds by the transformation

s : = λ v

. Equality (133) also holds for

λ = 0

since we have

D (P ∥ R_{0}) = D (P ∥ P) = 0

.

Proof of (23): For all

s \in (0, 1]

,

\begin{matrix} χ^{2} (P ∥ Q) & = \int \frac{{(p - q)}^{2}}{q} d μ \\ (134) & = \frac{1}{s^{2}} \int \frac{{[(s p + (1 - s) q) - q]}^{2}}{q} d μ \\ (135) & = \frac{1}{s^{2}} \int \frac{{(r_{1 - s} - q)}^{2}}{q} d μ \\ (136) & = \frac{1}{s^{2}} χ^{2} (R_{1 - s} ∥ Q), \end{matrix}

where (135) holds due to (122). From (136), it follows that for all

λ \in [0, 1]

,

\begin{matrix} \int_{0}^{λ} \frac{1}{s} χ^{2} (R_{1 - s} ∥ Q) d s = \int_{0}^{λ} s d s χ^{2} (P ∥ Q) = \frac{1}{2} λ^{2} χ^{2} (P ∥ Q) . \end{matrix}

(137)

5.2. Proof of Proposition 1

(a): Simple Proof of Pinsker’s Inequality: By [61] or [62] [(58)],

$χ^{2} (P ∥ Q) \geq {\begin{cases} {| P - Q |}^{2}, & if | P - Q | \in [0, 1], \\ \frac{| P - Q |}{2 - | P - Q |}, & if | P - Q | \in (1, 2] . \end{cases}$

(138)

We need the weaker inequality $χ^{2} (P ∥ Q) \geq {| P - Q |}^{2}$ , proved by the Cauchy–Schwarz inequality:

$\begin{matrix} (139) & χ^{2} (P ∥ Q) & = \int \frac{{(p - q)}^{2}}{q} d μ \int q d μ \\ (140) & \geq {(\int \frac{| p - q |}{\sqrt{q}} \cdot \sqrt{q} d μ)}^{2} \\ (141) & = {| P - Q |}^{2} . \end{matrix}$

By combining (24) and (139)–(141), it follows that

$\begin{matrix} (142) & \frac{1}{log e} D (P ∥ Q) & = \int_{0}^{1} χ^{2} (P ∥ (1 - s) P + s Q) \frac{d s}{s} \\ (143) & \geq \int_{0}^{1} | P - ((1 - s) P + s Q) |^{2} \frac{d s}{s} \\ (144) & = \int_{0}^{1} s {| P - Q |}^{2} d s \\ (145) & = \frac{1}{2} {| P - Q |}^{2} . \end{matrix}$
(b): Proof of (30) and its local tightness:

$\begin{matrix} (146) & \frac{1}{log e} D (P ∥ Q) & = \int_{0}^{1} χ^{2} (P ∥ (1 - s) P + s Q) \frac{d s}{s} \\ (147) & = \int_{0}^{1} (\int \frac{{[p - ((1 - s) p + s q)]}^{2}}{(1 - s) p + s q} d μ) \frac{d s}{s} \\ (148) & = \int_{0}^{1} \int \frac{s {(p - q)}^{2}}{(1 - s) p + s q} d μ d s \\ (149) & \leq \int_{0}^{1} \int s {(p - q)}^{2} (\frac{1 - s}{p} + \frac{s}{q}) d μ d s \\ (150) & = \int_{0}^{1} s^{2} d s \int \frac{{(p - q)}^{2}}{q} d μ + \int_{0}^{1} s (1 - s) d s \int \frac{{(p - q)}^{2}}{p} d μ \\ (151) & = \frac{1}{3} χ^{2} (P ∥ Q) + \frac{1}{6} χ^{2} (Q ∥ P), \end{matrix}$

where (146) is (24), and (149) holds due to Jensen’s inequality and the convexity of the hyperbola.
We next show the local tightness of inequality (30) by proving that (31) yields (32). Let ${P_{n}}$ be a sequence of probability measures, defined on a measurable space $(X, F)$ , and assume that ${P_{n}}$ converges to a probability measure P in the sense that (31) holds. In view of [16] [Theorem 7] (see also [15] [Section 4.F] and [63]), it follows that

$\begin{matrix} lim_{n \to \infty} D (P_{n} ∥ P) = lim_{n \to \infty} χ^{2} (P_{n} ∥ P) = 0, \end{matrix}$

(152)

and

$\begin{matrix} (153) & lim_{n \to \infty} \frac{D (P_{n} ∥ P)}{χ^{2} (P_{n} ∥ P)} = \frac{1}{2} log e, \\ (154) & lim_{n \to \infty} \frac{χ^{2} (P_{n} ∥ P)}{χ^{2} (P ∥ P_{n})} = 1, \end{matrix}$

which therefore yields (32).
(c): Proof of (33) and (34): The proof of (33) relies on (28) and the following lemma.
Lemma 1.
For all $s, θ \in (0, 1)$ ,

$\begin{matrix} \frac{D_{ϕ_{s}} (P ∥ Q)}{D_{ϕ_{θ}} (P ∥ Q)} \geq min \{\frac{1 - θ}{1 - s}, \frac{θ}{s}\} . \end{matrix}$

(155)

Proof.

$\begin{matrix} (156) & D_{ϕ_{s}} (P ∥ Q) & = \int \frac{{(p - q)}^{2}}{(1 - s) p + s q} d μ \\ (157) & = \int \frac{{(p - q)}^{2}}{(1 - θ) p + θ q} \frac{(1 - θ) p + θ q}{(1 - s) p + s q} d μ \\ (158) & \geq min \{\frac{1 - θ}{1 - s}, \frac{θ}{s}\} \int \frac{{(p - q)}^{2}}{(1 - θ) p + θ q} d μ \\ (159) & = min \{\frac{1 - θ}{1 - s}, \frac{θ}{s}\} D_{ϕ_{θ}} (P ∥ Q) . \end{matrix}$

□
From (28) and (155), for all $θ \in (0, 1)$ ,

$\begin{matrix} (160) & \frac{1}{log e} D (P ∥ Q) & = \int_{0}^{θ} s D_{ϕ_{s}} (P ∥ Q) d s + \int_{θ}^{1} s D_{ϕ_{s}} (P ∥ Q) d s \\ (161) & \geq \int_{0}^{θ} \frac{s (1 - θ)}{1 - s} \cdot D_{ϕ_{θ}} (P ∥ Q) d s + \int_{θ}^{1} θ D_{ϕ_{θ}} (P ∥ Q) d s \\ (162) & = [- θ + {log}_{e} (\frac{1}{1 - θ})] (1 - θ) D_{ϕ_{θ}} (P ∥ Q) + θ (1 - θ) D_{ϕ_{θ}} (P ∥ Q) \\ (163) & = (1 - θ) {log}_{e} (\frac{1}{1 - θ}) D_{ϕ_{θ}} (P ∥ Q) . \end{matrix}$

This proves (33). Furthermore, under the assumption in (31), for all $θ \in [0, 1]$ ,

$\begin{matrix} (164) & lim_{n \to \infty} \frac{D (P ∥ P_{n})}{D_{ϕ_{θ}} (P ∥ P_{n})} & = lim_{n \to \infty} \frac{D (P ∥ P_{n})}{χ^{2} (P ∥ P_{n})} lim_{n \to \infty} \frac{χ^{2} (P ∥ P_{n})}{D_{ϕ_{θ}} (P ∥ P_{n})} \\ (165) & = \frac{1}{2} log e \cdot \frac{2}{ϕ_{θ}^{″} (1)} \\ (166) & = \frac{1}{2} log e, \end{matrix}$

where (165) holds due to (153) and the local behavior of f-divergences [63], and (166) holds due to (17) which implies that $ϕ_{θ}^{″} (1) = 2$ for all $θ \in [0, 1]$ . This proves (34).
(d): Proof of (35): From (24), we get

$\begin{matrix} (167) & \frac{1}{log e} D (P ∥ Q) & = \int_{0}^{1} χ^{2} (P ∥ (1 - s) P + s Q) \frac{d s}{s} \\ (168) & = \int_{0}^{1} [χ^{2} (P ∥ (1 - s) P + s Q) - s^{2} χ^{2} (P ∥ Q)] \frac{d s}{s} + \int_{0}^{1} s d s χ^{2} (P ∥ Q) \\ (169) & = \int_{0}^{1} [χ^{2} (P ∥ (1 - s) P + s Q) - s^{2} χ^{2} (P ∥ Q)] \frac{d s}{s} + \frac{1}{2} χ^{2} (P ∥ Q) . \end{matrix}$

Referring to the integrand of the first term on the right side of (169), for all $s \in (0, 1]$ ,

$\begin{array}{l} \frac{1}{s} [χ^{2} (P ∥ (1 - s) P + s Q) - s^{2} χ^{2} (P ∥ Q)] \\ (170) & = s \int {(p - q)}^{2} [\frac{1}{(1 - s) p + s q} - \frac{1}{q}] d μ \\ (171) & = s (1 - s) \int \frac{{(q - p)}^{3}}{q [(1 - s) p + s q]} d μ \\ (172) & = s (1 - s) \int | q - p | \cdot \underset{\leq \frac{1}{s} 1 {q \geq p}}{\underset{︸}{\frac{| q - p |}{q} \cdot \frac{q - p}{p + s (q - p)}}} d μ \\ (173) & \leq (1 - s) \int (q - p) 1 {q \geq p} d μ \\ (174) & = \frac{1}{2} (1 - s) | P - Q |, \end{array}$

where the last equality holds since the equality $\int (q - p) d μ = 0$ implies that

$\begin{matrix} (175) & \int (q - p) 1 {q \geq p} d μ & = \int (p - q) 1 {p \geq q} d μ \\ (176) & = \frac{1}{2} \int | p - q | d μ = \frac{1}{2} | P - Q | . \end{matrix}$

From (170)–(174), an upper bound on the right side of (169) results. This gives

$\begin{matrix} (177) & \frac{1}{log e} D (P ∥ Q) & \leq \frac{1}{2} \int_{0}^{1} (1 - s) d s | P - Q | + \frac{1}{2} χ^{2} (P ∥ Q) \\ (178) & = \frac{1}{4} | P - Q | + \frac{1}{2} χ^{2} (P ∥ Q) . \end{matrix}$

It should be noted that [15] [Theorem 2(a)] shows that inequality (35) is tight. To that end, let $ε \in (0, 1)$ , and define probability measures $P_{ε}$ and $Q_{ε}$ on the set $A = {0, 1}$ with $P_{ε} (1) = ε^{2}$ and $Q_{ε} (1) = ε$ . Then,

$\begin{matrix} lim_{ε ↓ 0} \frac{\frac{1}{log e} D (P_{ε} ∥ Q_{ε})}{\frac{1}{4} | P_{ε} - Q_{ε} | + \frac{1}{2} χ^{2} (P_{ε} ∥ Q_{ε})} = 1 . \end{matrix}$

(179)

5.3. Proof of Theorem 2

We first prove Item (a) in Theorem 2. In view of the Hammersley–Chapman–Robbins lower bound on the

χ^{2}

divergence, for all

λ \in [0, 1]

\begin{matrix} χ^{2} (P ∥ (1 - λ) P + λ Q) \geq \frac{{(E [X] - E [Z_{λ}])}^{2}}{Var (Z_{λ})}, \end{matrix}

(180)

where

X \sim P

,

Y \sim Q

and

Z_{λ} \sim R_{λ} : = (1 - λ) P + λ Q

is defined by

Z_{λ} : = {\begin{cases} X, & with probability 1 - λ, \\ Y, & with probability λ . \end{cases}

(181)

For

λ \in [0, 1]

,

\begin{matrix} E [Z_{λ}] = (1 - λ) m_{P} + λ m_{Q}, \end{matrix}

(182)

and it can be verified that

\begin{matrix} Var (Z_{λ}) = (1 - λ) σ_{P}^{2} + λ σ_{Q}^{2} + λ (1 - λ) {(m_{P} - m_{Q})}^{2} . \end{matrix}

(183)

We now rely on identity (24)

\begin{matrix} \frac{1}{log e} D (P ∥ Q) & = \int_{0}^{1} χ^{2} (P ∥ (1 - λ) P + λ Q) \frac{d λ}{λ} \end{matrix}

(184)

to get a lower bound on the relative entropy. Combining (180), (183) and (184) yields

\begin{matrix} \frac{1}{log e} D (P ∥ Q) \geq {(m_{P} - m_{Q})}^{2} \int_{0}^{1} \frac{λ}{(1 - λ) σ_{P}^{2} + λ σ_{Q}^{2} + λ (1 - λ) {(m_{P} - m_{Q})}^{2}} d λ . \end{matrix}

(185)

From (43) and (44), we get

\begin{matrix} \int_{0}^{1} \frac{λ}{(1 - λ) σ_{P}^{2} + λ σ_{Q}^{2} + λ (1 - λ) {(m_{P} - m_{Q})}^{2}} d λ = \int_{0}^{1} \frac{λ}{(α - a λ) (β + a λ)} d λ, \end{matrix}

(186)

where

\begin{matrix} α & : = \sqrt{σ_{P}^{2} + \frac{b^{2}}{4 a^{2}}} + \frac{b}{2 a}, \end{matrix}

(187)

\begin{matrix} β & : = \sqrt{σ_{P}^{2} + \frac{b^{2}}{4 a^{2}}} - \frac{b}{2 a} . \end{matrix}

(188)

By using the partial fraction decomposition of the integrand on the right side of (186), we get (after multiplying both sides of (185) by

log e

)

\begin{matrix} (189) & D (P ∥ Q) & \geq \frac{{(m_{P} - m_{Q})}^{2}}{a^{2}} [\frac{α}{α + β} log (\frac{α}{α - a}) + \frac{β}{α + β} log (\frac{β}{β + a})] \\ (190) & = \frac{α}{α + β} log (\frac{α}{α - a}) + \frac{β}{α + β} log (\frac{β}{β + a}) \\ (191) & = d (\frac{α}{α + β} ∥\frac{α - a}{α + β}), \end{matrix}

where (189) holds by integration since

α - a λ

and

β + a λ

are both non-negative for all

λ \in [0, 1]

. To verify the latter claim, it should be noted that (43) and the assumption that

m_{P} \neq m_{Q}

imply that

a \neq 0

. Since

α, β > 0

, it follows that, for all

λ \in [0, 1]

, either

α - a λ > 0

or

β + a λ > 0

(if

a < 0

, then the former is positive, and, if

a > 0

, then the latter is positive). By comparing the denominators of both integrands on the left and right sides of (186), it follows that

(α - a λ) (β + a λ) \geq 0

for all

λ \in [0, 1]

. Since the product of

α - a λ

and

β + a λ

is non-negative and at least one of these terms is positive, it follows that

α - a λ

and

β + a λ

are both non-negative for all

λ \in [0, 1]

. Finally, (190) follows from (43).

If

m_{P} - m_{Q} \to 0

and

σ_{P} \neq σ_{Q}

, then it follows from (43) and (44) that

a \to 0

and

b \to σ_{P}^{2} - σ_{Q}^{2} \neq 0

. Hence, from (187) and (188),

α \geq |\frac{b}{a}| \to \infty

and

β \to 0

, which implies that the lower bound on

D (P ∥ Q)

in (191) tends to zero.

Letting

r : = \frac{α}{α + β}

and

s : = \frac{α - a}{α + β}

, we obtain that the lower bound on

D (P ∥ Q)

in (40) holds. This bound is consistent with the expressions of r and s in (41) and (42) since, from (45), (187) and (188),

\begin{matrix} r & = \frac{α}{α + β} = \frac{v + \frac{b}{2 a}}{2 v} = \frac{1}{2} + \frac{b}{4 a v}, \end{matrix}

(192)

\begin{matrix} s & = \frac{α - a}{α + β} = r - \frac{a}{α + β} = r - \frac{a}{2 v} . \end{matrix}

(193)

It should be noted that

r, s \in [0, 1]

. First, from (187) and (188),

α

and

β

are positive if

σ_{P} \neq 0

, which yields

r = \frac{α}{α + β} \in (0, 1)

. We next show that

s \in [0, 1]

. Recall that

α - a λ

and

β + a λ

are both non-negative for all

λ \in [0, 1]

. Setting

λ = 1

yields

α \geq a

, which (from (193)) implies that

s \geq 0

. Furthermore, from (193) and the positivity of

α + β

, it follows that

s \leq 1

if and only if

β \geq - a

. The latter holds since

β + a λ \geq 0

for all

λ \in [0, 1]

(in particular, for

λ = 1

). If

σ_{P} = 0

, then it follows from (41)–(45) that

v = \frac{b}{2 | a |}

,

b = a^{2} + σ_{Q}^{2}

, and (recall that

a \neq 0

)

(i): If $a > 0$ , then $v = \frac{b}{2 a}$ implies that $r = \frac{1}{2} + \frac{b}{4 a v} = 1$ , and $s = r - \frac{a}{2 v} = 1 - \frac{a^{2}}{b} = \frac{σ_{Q}^{2}}{σ_{Q}^{2} + a^{2}} \in [0, 1]$ ;
(ii): If $a < 0$ , then $v = - \frac{b}{2 a}$ implies that $r = 0$ , and $s = r - \frac{a}{2 v} = \frac{a^{2}}{b} = \frac{a^{2}}{a^{2} + σ_{Q}^{2}} \in [0, 1]$ .

We next prove Item (b) in Theorem 2 (i.e., the achievability of the lower bound in (40)). To that end, we provide a technical lemma, which can be verified by the reader.

Lemma 2.

Let

r, s

be given in (41)–(45), and let

u_{1, 2}

be given in (47). Then,

\begin{matrix} (s - r) (u_{1} - u_{2}) = m_{Q} - m_{P}, \end{matrix}

(194)

\begin{matrix} u_{1} + u_{2} = m_{P} + m_{Q} + \frac{σ_{Q}^{2} - σ_{P}^{2}}{m_{Q} - m_{P}} . \end{matrix}

(195)

Let

X \sim P

and

Y \sim Q

be defined on a set

U = {u_{1}, u_{2}}

(for the moment, the values of

u_{1}

and

u_{2}

are not yet specified) with

P [X = u_{1}] = r

,

P [X = u_{2}] = 1 - r

,

Q [Y = u_{1}] = s

, and

Q [Y = u_{2}] = 1 - s

. We now calculate

u_{1}

and

u_{2}

such that

E [X] = m_{P}

and

Var (X) = σ_{P}^{2}

. This is equivalent to

\begin{array}{l} (196) & r u_{1} + (1 - r) u_{2} = m_{P}, \\ (197) & r u_{1}^{2} + (1 - r) u_{2}^{2} = m_{P}^{2} + σ_{P}^{2} . \end{array}

Substituting (196) into the right side of (197) gives

\begin{matrix} r u_{1}^{2} + (1 - r) u_{2}^{2} = {[r u_{1} + (1 - r) u_{2}]}^{2} + σ_{P}^{2}, \end{matrix}

(198)

which, by rearranging terms, also gives

\begin{matrix} u_{1} - u_{2} = \pm \sqrt{\frac{σ_{P}^{2}}{r (1 - r)}} . \end{matrix}

(199)

Solving simultaneously (196) and (199) gives

\begin{matrix} u_{1} = m_{P} \pm \sqrt{\frac{(1 - r) σ_{P}^{2}}{r}}, \end{matrix}

(200)

\begin{matrix} u_{2} = m_{P} \mp \sqrt{\frac{r σ_{P}^{2}}{1 - r}} . \end{matrix}

(201)

We next verify that, by setting

u_{1, 2}

as in (47), one also gets (as desired) that

E [Y] = m_{Q}

and

Var (Y) = σ_{Q}^{2}

. From Lemma 2, and, from (196) and (197), we have

\begin{matrix} (202) & E [Y] & = s u_{1} + (1 - s) u_{2} \\ (203) & = (r u_{1} + (1 - r) u_{2}) + (s - r) (u_{1} - u_{2}) \\ (204) & = m_{P} + (s - r) (u_{1} - u_{2}) = m_{Q}, \\ (205) & E [Y^{2}] & = s u_{1}^{2} + (1 - s) u_{2}^{2} \\ (206) & = r u_{1}^{2} + (1 - r) u_{2}^{2} + (s - r) (u_{1}^{2} - u_{2}^{2}) \\ (207) & = E [X^{2}] + (s - r) (u_{1} - u_{2}) (u_{1} + u_{2}) \\ (208) & = m_{P}^{2} + σ_{P}^{2} + (m_{Q} - m_{P}) (m_{P} + m_{Q} + \frac{σ_{Q}^{2} - σ_{P}^{2}}{m_{Q} - m_{P}}) \\ (209) & = m_{Q}^{2} + σ_{Q}^{2} . \end{matrix}

By combining (204) and (209), we obtain

Var (Y) = σ_{Q}^{2}

. Hence, the probability mass functions P and Q defined on

U = {u_{1}, u_{2}}

(with

u_{1}

and

u_{2}

in (47)) such that

\begin{matrix} P (u_{1}) = 1 - P (u_{2}) = r, Q (u_{1}) = 1 - Q (u_{2}) = s \end{matrix}

(210)

satisfy the equality constraints in (39), while also achieving the lower bound on

D (P ∥ Q)

that is equal to

d (r ∥ s)

. It can be also verified that the second option where

\begin{matrix} u_{1} = m_{P} - \sqrt{\frac{(1 - r) σ_{P}^{2}}{r}}, u_{2} = m_{P} + \sqrt{\frac{r σ_{P}^{2}}{1 - r}} \end{matrix}

(211)

does not yield the satisfiability of the conditions

E [Y] = m_{Q}

and

Var (Y) = σ_{Q}^{2}

, so there is only a unique pair of probability measures P and Q, defined on a two-element set that achieves the lower bound in (40) under the equality constraints in (39).

We finally prove Item (c) in Theorem 2. Let

m \in R, σ_{P}^{2}

, and

σ_{Q}^{2}

be selected arbitrarily such that

σ_{Q}^{2} \geq σ_{P}^{2}

. We construct probability measures

P_{ε}

and

Q_{ε}

, depending on a free parameter

ε

, with means

m_{P} = m_{Q} : = m

and variances

σ_{P}^{2}

and

σ_{Q}^{2}

, respectively (means and variances are independent of

ε

), and which are defined on a three-element set

U : = {u_{1}, u_{2}, u_{3}}

as follows:

\begin{matrix} P_{ε} (u_{1}) = r, P_{ε} (u_{2}) = 1 - r, P_{ε} (u_{3}) = 0, \end{matrix}

(212)

\begin{matrix} Q_{ε} (u_{1}) = s, Q_{ε} (u_{2}) = 1 - s - ε, Q_{ε} (u_{3}) = ε, \end{matrix}

(213)

with

ε > 0

. We aim to set the parameters

r, s, u_{1}, u_{2}

and

u_{3}

(as a function of

m, σ_{P}, σ_{Q}

and

ε

) such that

\begin{matrix} lim_{ε \to 0^{+}} D (P_{ε} ∥ Q_{ε}) = 0 . \end{matrix}

(214)

Proving (214) yields (48), while it also follows that the infimum on the left side of (48) can be restricted to probability measures which are defined on a three-element set.

In view of the constraints on the means and variances in (39), with equal means m, we get the following set of equations from (212) and (213):

\begin{matrix} \{\begin{matrix} r u_{1} + (1 - r) u_{2} = m, \\ s u_{1} + (1 - s - ε) u_{2} + ε u_{3} = m, \\ r u_{1}^{2} + (1 - r) u_{2}^{2} = m^{2} + σ_{P}^{2}, \\ s u_{1}^{2} + (1 - s - ε) u_{2}^{2} + ε u_{3}^{2} = m^{2} + σ_{Q}^{2} . \end{matrix} \end{matrix}

(215)

The first and second equations in (215) refer to the equal means under P and Q, and the third and fourth equations in (215) refer to the second moments in (39). Furthermore, in view of (212) and (213), the relative entropy is given by

\begin{matrix} D (P_{ε} ∥ Q_{ε}) = r log \frac{r}{s} + (1 - r) log \frac{1 - r}{1 - s - ε} . \end{matrix}

(216)

Subtracting the square of the first equation in (215) from its third equation gives the equivalent set of equations

\begin{matrix} \{\begin{matrix} r u_{1} + (1 - r) u_{2} = m, \\ s u_{1} + (1 - s - ε) u_{2} + ε u_{3} = m, \\ r (1 - r) {(u_{1} - u_{2})}^{2} = σ_{P}^{2}, \\ s u_{1}^{2} + (1 - s - ε) u_{2}^{2} + ε u_{3}^{2} = m^{2} + σ_{Q}^{2} . \end{matrix} \end{matrix}

(217)

We next select

u_{1}

and

u_{2}

such that

u_{1} - u_{2} : = 2 σ_{P}

. Then, the third equation in (217) gives

r (1 - r) = \frac{1}{4}

, so

r = \frac{1}{2}

. Furthermore, the first equation in (217) gives

\begin{matrix} u_{1} = m + σ_{P}, \end{matrix}

(218)

\begin{matrix} u_{2} = m - σ_{P} . \end{matrix}

(219)

Since r,

u_{1}

, and

u_{2}

are independent of

ε

, so is the probability measure

P_{ε} : = P

. Combining the second equation in (217) with (218) and (219) gives

\begin{matrix} u_{3} = m - (1 + \frac{2 s - 1}{ε}) σ_{P} . \end{matrix}

(220)

Substituting (218)–(220) into the fourth equation of (217) gives a quadratic equation for s, whose selected solution (such that s and

r = \frac{1}{2}

be close for small

ϵ > 0

) is equal to

\begin{matrix} s = \frac{1}{2} [1 - ε + \sqrt{(\frac{σ_{Q}^{2}}{σ_{P}^{2}} - 1 + ε) ε}] . \end{matrix}

(221)

Hence,

s = \frac{1}{2} + O (\sqrt{ε})

, which implies that

s \in (0, 1 - ε)

for sufficiently small

ε > 0

(as it is required in (213)). In view of (216), it also follows that

D (P ∥ Q_{ε})

vanishes as we let

ε

tend to zero.

We finally outline an alternative proof, which refers to the case of equal means with arbitrarily selected

σ_{P}^{2}

and

σ_{Q}^{2}

. Let

(σ_{P}^{2}, σ_{Q}^{2}) \in {(0, \infty)}^{2}

. We next construct a sequence of pairs of probability measures

{(P_{n}, Q_{n})}

with zero mean and respective variances

(σ_{P}^{2}, σ_{Q}^{2})

for which

D (P_{n} ∥ Q_{n}) \to 0

as

n \to \infty

(without any loss of generality, one can assume that the equal means are equal to zero). We start by assuming

(σ_{P}^{2}, σ_{Q}^{2}) \in {(1, \infty)}^{2}

. Let

\begin{matrix} μ_{n} : = \sqrt{1 + n (σ_{Q}^{2} - 1)}, \end{matrix}

(222)

and define a sequence of quaternary real-valued random variables with probability mass functions

Q_{n} (a) : = {\begin{cases} \frac{1}{2} - \frac{1}{2 n} & a = \pm 1, \\ \frac{1}{2 n} & a = \pm μ_{n} . \end{cases}

(223)

It can be verified that, for all

n \in N

,

Q_{n}

has zero mean and variance

σ_{Q}^{2}

. Furthermore, let

P_{n} (a) : = {\begin{cases} \frac{1}{2} - \frac{ξ}{2 n} & a = \pm 1, \\ \frac{ξ}{2 n} & a = \pm μ_{n}, \end{cases}

(224)

with

\begin{matrix} ξ : = \frac{σ_{P}^{2} - 1}{σ_{Q}^{2} - 1} . \end{matrix}

(225)

If

ξ > 1

, for

n = 1, \dots, ⌈ ξ ⌉

, we choose

P_{n}

arbitrarily with mean 0 and variance

σ_{P}^{2}

. Then,

\begin{matrix} Var (P_{n}) = 1 - \frac{ξ}{n} + \frac{ξ}{n} μ_{n}^{2} = σ_{P}^{2}, \end{matrix}

(226)

\begin{matrix} D (P_{n} ∥ Q_{n}) = d (\frac{ξ}{n} ∥ \frac{1}{n}) \to 0 . \end{matrix}

(227)

Next, suppose

min {σ_{P}^{2}, σ_{Q}^{2}} : = σ^{2} < 1

, then construct

P_{n}^{'}

and

Q_{n}^{'}

as before with variances

\frac{2 σ_{P}^{2}}{σ^{2}} > 1

and

\frac{2 σ_{Q}^{2}}{σ^{2}} > 1

, respectively. If

P_{n}

and

Q_{n}

denote the random variables

P_{n}^{'}

and

Q_{n}^{'}

scaled by a factor of

\frac{σ}{\sqrt{2}}

, then their variances are

σ_{P}^{2}

,

σ_{Q}^{2}

, respectively, and

D (P_{n} ∥ Q_{n}) = D (P_{n}^{'} ∥ Q_{n}^{'}) \to 0

as we let

n \to \infty

.

To conclude, it should be noted that the sequences of probability measures in the latter proof are defined on a four-element set. Recall that, in the earlier proof, specialized to the case of (equal means with)

σ_{P}^{2} \leq σ_{Q}^{2}

, the introduced probability measures are defined on a three-element set, and the reference probability measure P is fixed while referring to an equiprobable binary random variable.

5.4. Proof of Theorem 3

We first prove (52). Differentiating both sides of (22) gives that, for all

λ \in (0, 1]

,

\begin{matrix} (228) & F^{'} (λ) & = \frac{1}{λ} χ^{2} (P ∥ R_{λ}) log e \\ (229) & \geq \frac{1}{λ} [exp (D (P ∥ R_{λ})) - 1] log e \\ (230) & = \frac{1}{λ} [exp (F (λ)) - 1] log e, \end{matrix}

where (228) holds due to (21), (22) and (50); (229) holds by (16) and (230) is due to (21) and (50). This gives (52).

We next prove (53), and the conclusion which appears after it. In view of [16] [Theorem 8], applied to

f (t) : = - log t

for all

t > 0

, we get (it should be noted that, by the definition of F in (50), the result in [16] [(195)–(196)] is used here by swapping P and Q)

\begin{matrix} lim_{λ \to 0^{+}} \frac{F (λ)}{λ^{2}} = \frac{1}{2} χ^{2} (Q ∥ P) log e . \end{matrix}

(231)

Since

lim_{λ \to 0^{+}} F (λ) = 0

, it follows by L’Hôpital’s rule that

\begin{matrix} lim_{λ \to 0^{+}} \frac{F^{'} (λ)}{λ} = 2 lim_{λ \to 0^{+}} \frac{F (λ)}{λ^{2}} = χ^{2} (Q ∥ P) log e, \end{matrix}

(232)

which gives (53). A comparison of the limit in (53) with a lower bound which follows from (52) gives

\begin{matrix} (233) & lim_{λ \to 0^{+}} \frac{F^{'} (λ)}{λ} & \geq lim_{λ \to 0^{+}} \frac{1}{λ^{2}} [exp (F (λ)) - 1] log e \\ (234) & = lim_{λ \to 0^{+}} \frac{F (λ)}{λ^{2}} lim_{λ \to 0^{+}} \frac{exp (F (λ)) - 1}{F (λ)} \cdot log e \\ (235) & = lim_{λ \to 0^{+}} \frac{F (λ)}{λ^{2}} lim_{u \to 0} \frac{e^{u} - 1}{u} \\ (236) & = \frac{1}{2} χ^{2} (Q ∥ P) log e, \end{matrix}

where (236) relies on (231). Hence, the limit in (53) is twice as large as its lower bound on the right side of (236). This proves the conclusion which comes right after (53).

We finally prove the known result in (51), by showing an alternative proof which is based on (52). The function F is non-negative on

[0, 1]

, and it is strictly positive on

(0, 1]

if

P \neq Q

. Let

P \neq Q

(otherwise, (51) is trivial). Rearranging terms in (52) and integrating both sides over the interval

[λ, 1]

, for

λ \in (0, 1]

, gives that

\begin{matrix} (237) & \int_{λ}^{1} \frac{F^{'} (t)}{exp (F (t)) - 1} d t & \geq \int_{λ}^{1} \frac{d t}{t} log e \\ (238) & = log \frac{1}{λ}, \forall λ \in (0, 1] . \end{matrix}

The left side of (237) satisfies

\begin{matrix} (239) & \int_{λ}^{1} \frac{F^{'} (t)}{exp (F (t)) - 1} d t & = \int_{λ}^{1} \frac{F^{'} (t) exp (- F (t))}{1 - exp (- F (t))} d t \\ (240) & = \int_{λ}^{1} \frac{d}{d t} \{log (1 - exp (- F (t)))\} d t \\ (241) & = log (\frac{1 - exp (- D (P ∥ Q))}{1 - exp (- F (λ))}), \end{matrix}

where (241) holds since

F (1) = D (P ∥ Q)

(see (50)). Combining (237)–(241) gives

\begin{matrix} \frac{1 - exp (- D (P ∥ Q))}{1 - exp (- F (λ))} \geq \frac{1}{λ}, \forall λ \in (0, 1], \end{matrix}

(242)

which, due to the non-negativity of F, gives the right side inequality in (51) after rearrangement of terms in (242).

5.5. Proof of Theorem 4

Lemma 3.

Let

f_{0} : (0, \infty) \to R

be a convex function with

f_{0} (1) = 0

, and let

{f_{k} (\cdot)}_{k = 0}^{\infty}

be defined as in (58). Then,

{f_{k} (\cdot)}_{k = 0}^{\infty}

is a sequence of convex functions on

(0, \infty)

, and

\begin{matrix} f_{k} (x) \geq f_{k + 1} (x), \forall x > 0, k \in {0, 1, \dots} . \end{matrix}

(243)

Proof.

We prove the convexity of

{f_{k} (\cdot)}

on

(0, \infty)

by induction. Suppose that

f_{k} (\cdot)

is a convex function with

f_{k} (1) = 0

for a fixed integer

k \geq 0

. The recursion in (58) yields

f_{k + 1} (1) = 0

and, by the change of integration variable

s : = (1 - x) s^{'}

,

\begin{matrix} f_{k + 1} (x) & = \int_{0}^{1} f_{k} (s^{'} x - s^{'} + 1) \frac{d s^{'}}{s^{'}}, x > 0 . \end{matrix}

(244)

Consequently, for

t \in (0, 1)

and

x \neq y

with

x, y > 0

, applying (244) gives

\begin{matrix} (245) & f_{k + 1} ((1 - t) x + t y) & = \int_{0}^{1} f_{k} (s^{'} [(1 - t) x + t y] - s^{'} + 1) \frac{d s^{'}}{s^{'}} \\ (246) & = \int_{0}^{1} f_{k} ((1 - t) (s^{'} x - s^{'} + 1) + t (s^{'} y - s^{'} + 1)) \frac{d s^{'}}{s^{'}} \\ (247) & \leq (1 - t) \int_{0}^{1} f_{k} (s^{'} x - s^{'} + 1) \frac{d s^{'}}{s^{'}} + t \int_{0}^{1} f_{k} (s^{'} y - s^{'} + 1) \frac{d s^{'}}{s^{'}} \\ (248) & = (1 - t) f_{k + 1} (x) + t f_{k + 1} (y), \end{matrix}

where (247) holds since

f_{k} (\cdot)

is convex on

(0, \infty)

(by assumption). Hence, from (245)–(248),

f_{k + 1} (\cdot)

is also convex on

(0, \infty)

with

f_{k + 1} (1) = 0

. By mathematical induction and our assumptions on

f_{0}

, it follows that

{f_{k} (\cdot)}_{k = 0}^{\infty}

is a sequence of convex functions on

(0, \infty)

which vanish at 1.

We next prove (243). For all

x, y > 0

and

k \in {0, 1, \dots}

,

\begin{matrix} (249) & f_{k + 1} (y) & \geq f_{k + 1} (x) + f_{k + 1}^{'} (x) (y - x) \\ (250) & = f_{k + 1} (x) + \frac{f_{k} (x)}{x - 1} (y - x), \end{matrix}

where (249) holds since

f_{k} (\cdot)

is convex on

(0, \infty)

, and (250) relies on the recursive equation in (58). Substituting

y = 1

into (249)–(250), and using the equality

f_{k + 1} (1) = 0

, gives (243). □

We next prove Theorem 4. From Lemma 3, it follows that

D_{f_{k}} (P ∥ Q)

is an f-divergence for all integers

k \geq 0

, and the non-negative sequence

{D_{f_{k}} {(P ∥ Q)}}_{k = 0}^{\infty}

is monotonically non-increasing. From (21) and (58), it also follows that, for all

λ \in [0, 1]

and integer

k \in {0, 1, \dots}

,

\begin{matrix} (251) & D_{f_{k + 1}} (R_{λ} ∥ P) & = \int p f_{k + 1} (\frac{r_{λ}}{p}) d μ \\ (252) & = \int p \int_{0}^{(p - q) λ / p} f_{k} (1 - s) \frac{d s}{s} d μ \\ (253) & = \int p \int_{0}^{λ} f_{k} (1 + \frac{(q - p) s^{'}}{p}) \frac{d s^{'}}{s^{'}} d μ \\ (254) & = \int_{0}^{λ} \int p f_{k} (\frac{r_{s^{'}}}{p}) d μ \frac{d s^{'}}{s^{'}} \\ (255) & = \int_{0}^{λ} D_{f_{k}} (R_{s^{'}} ∥ P) \frac{d s^{'}}{s^{'}}, \end{matrix}

where the substitution

s : = \frac{(p - q) s^{'}}{p}

is invoked in (253), and then (254) holds since

\frac{r_{s^{'}}}{p} = 1 + \frac{(q - p) s^{'}}{p}

for

s^{'} \in [0, 1]

(this follows from (21)) and by interchanging the order of the integrations.

5.6. Proof of Corollary 5

Combining (60) and (61) yields (58); furthermore,

f_{0} : (0, \infty) \to R

, given by

f_{0} (x) = \frac{1}{x} - 1

for all

x > 0

, is convex on

(0, \infty)

with

f_{0} (1) = 0

. Hence, Theorem 4 holds for the selected functions

{f_{k} (\cdot)}_{k = 0}^{\infty}

in (61), which therefore are all convex on

(0, \infty)

and vanish at 1. This proves that (59) holds for all

λ \in [0, 1]

and

k \in {0, 1, \dots}

. Since

f_{0} (x) = \frac{1}{x} - 1

and

f_{1} (x) = - {log}_{e} (x)

for all

x > 0

(see (60) and (61)), then, for every pair of probability measures P and Q:

\begin{matrix} D_{f_{0}} (P ∥ Q) = χ^{2} (Q ∥ P), D_{f_{1}} (P ∥ Q) = \frac{1}{log e} D (Q ∥ P) . \end{matrix}

(256)

Finally, combining (59), for

k = 0

, together with (256), gives (22) as a special case.

5.7. Proof of Theorem 5 and Corollary 6

For an arbitrary measurable set

E \subseteq X

, we have from (62)

\begin{matrix} μ_{C} (E) = \int_{E} \frac{1_{C} (x)}{μ (C)} d μ (x), \end{matrix}

(257)

where

1_{C} : X \to {0, 1}

is the indicator function of

C \subseteq X

, i.e.,

1_{C} (x) : = 1 {x \in C}

for

x \in X

. Hence,

\begin{matrix} \frac{d μ_{C}}{d μ} (x) = \frac{1_{C} (x)}{μ (C)}, \forall x \in X, \end{matrix}

(258)

and

\begin{matrix} (259) & D (μ_{C} ∥ μ) & = \int_{X} f (\frac{d μ_{C}}{d μ}) d μ \\ (260) & = \int_{C} f (\frac{1}{μ (C)}) d μ (x) + \int_{X \ C} f (0) d μ (x) \\ (261) & = μ (C) f (\frac{1}{μ (C)}) + μ (X \ C) f (0) \\ (262) & = \tilde{f} (μ (C)) + (1 - μ (C)) f (0), \end{matrix}

where the last equality holds by the definition of

\tilde{f}

in (63). This proves Theorem 5. Corollary 6 is next proved by first proving (67) for the Rényi divergence. For all

α \in (0, 1) \cup (1, \infty)

,

\begin{matrix} (263) & D_{α} (μ_{C} ∥ μ) & = \frac{1}{α - 1} log \int_{X} {(\frac{d μ_{C}}{d μ})}^{α} d μ \\ (264) & = \frac{1}{α - 1} log \int_{C} {(\frac{1}{μ (C)})}^{α} d μ \\ (265) & = \frac{1}{α - 1} log ({(\frac{1}{μ (C)})}^{α} μ (C)) \\ (266) & = log \frac{1}{μ (C)} . \end{matrix}

The justification of (67) for

α = 1

is due to the continuous extension of the order-

α

Rényi divergence at

α = 1

, which gives the relative entropy (see (13)). Equality (65) is obtained from (67) at

α = 1

. Finally, (66) is obtained by combining (15) and (67) with

α = 2

.

5.8. Proof of Theorem 6

Equation (100) is an equivalent form of (27). From (91) and (100), for all

α \in [0, 1]

,

\begin{matrix} (267) & \frac{1}{log e} S_{α} (P ∥ Q) & = α \frac{1}{log e} K_{α} (P ∥ Q) + (1 - α) \frac{1}{log e} K_{1 - α} (Q ∥ P) \\ (268) & = α \int_{0}^{α} s D_{ϕ_{s}} (P ∥ Q) d s + (1 - α) \int_{0}^{1 - α} s D_{ϕ_{s}} (Q ∥ P) d s \\ (269) & = α \int_{0}^{α} s D_{ϕ_{s}} (P ∥ Q) d s + (1 - α) \int_{α}^{1} (1 - s) D_{ϕ_{1 - s}} (Q ∥ P) d s . \end{matrix}

Regarding the integrand of the second term in (269), in view of (18), for all

s \in (0, 1)

\begin{matrix} (270) & D_{ϕ_{1 - s}} (Q ∥ P) & = \frac{1}{{(1 - s)}^{2}} \cdot χ^{2} (Q ∥ (1 - s) P + s Q) \\ (271) & = \frac{1}{s^{2}} \cdot χ^{2} (P ∥ (1 - s) P + s Q) \\ (272) & = D_{ϕ_{s}} (P ∥ Q), \end{matrix}

where (271) readily follows from (9). Since we also have

D_{ϕ_{1}} (P ∥ Q) = χ^{2} (P ∥ Q) = D_{ϕ_{0}} (Q ∥ P)

(see (18)), it follows that

\begin{matrix} D_{ϕ_{1 - s}} (Q ∥ P) = D_{ϕ_{s}} (P ∥ Q), s \in [0, 1] . \end{matrix}

(273)

By using this identity, we get from (269) that, for all

α \in [0, 1]

\begin{matrix} (274) & \frac{1}{log e} S_{α} (P ∥ Q) & = α \int_{0}^{α} s D_{ϕ_{s}} (P ∥ Q) d s + (1 - α) \int_{α}^{1} (1 - s) D_{ϕ_{s}} (P ∥ Q) d s \\ (275) & = \int_{0}^{1} g_{α} (s) D_{ϕ_{s}} (P ∥ Q) d s, \end{matrix}

where the function

g_{α} : [0, 1] \to R

is defined in (102). This proves the integral identity (101).

The lower bounds in (103) and (104) hold since, if

f : (0, \infty) \to R

is convex, continuously twice differentiable and strictly convex at 1, then

\begin{matrix} μ_{χ^{2}} (Q_{X}, W_{Y | X}) \leq μ_{f} (Q_{X}, W_{Y | X}), \end{matrix}

(276)

(see, e.g., [46] [Proposition II.6.5] and [50] [Theorem 2]). Hence, this holds in particular for the f-divergences in (95) and (96) (since the required properties are satisfied by the parametric functions in (97) and (98), respectively). We next prove the upper bound on the contraction coefficients in (103) and (104) by relying on (100) and (101), respectively. In the setting of Definition 7, if

P_{X} \neq Q_{X}

, then it follows from (100) that for

α \in (0, 1]

,

\begin{matrix} (277) & \frac{K_{α} (P_{Y} ∥ Q_{Y})}{K_{α} (P_{X} ∥ Q_{X})} & = \frac{\int_{0}^{α} s D_{ϕ_{s}} (P_{Y} ∥ Q_{Y}) d s}{\int_{0}^{α} s D_{ϕ_{s}} (P_{X} ∥ Q_{X}) d s} \\ (278) & \leq \frac{\int_{0}^{α} s μ_{ϕ_{s}} (Q_{X}, W_{Y | X}) D_{ϕ_{s}} (P_{X} ∥ Q_{X}) d s}{\int_{0}^{α} s D_{ϕ_{s}} (P_{X} ∥ Q_{X}) d s} \\ (279) & \leq sup_{s \in (0, α]} μ_{ϕ_{s}} (Q_{X}, W_{Y | X}) . \end{matrix}

Finally, taking the supremum of the left-hand side of (277) over all probability measures

P_{X}

such that

0 < K_{α} (P_{X} ∥ Q_{X}) < \infty

gives the upper bound on

μ_{k_{α}} (Q_{X}, W_{Y | X})

in (103). The proof of the upper bound on

μ_{s_{α}} (Q_{X}, W_{Y | X})

, for all

α \in [0, 1]

, follows similarly from (101), since the function

g_{α} (\cdot)

as defined in (102) is positive over the interval

(0, 1)

.

5.9. Proof of Corollary 7

The upper bounds in (106) and (107) rely on those in (103) and (104), respectively, by showing that

\begin{matrix} sup_{s \in (0, 1]} μ_{ϕ_{s}} (Q_{X}, W_{Y | X}) \leq μ_{χ^{2}} (W_{Y | X}) . \end{matrix}

(280)

Inequality (280) is obtained as follows, similarly to the concept of the proof of [51] [Remark 3.8]. For all

s \in (0, 1]

and

P_{X} \neq Q_{X}

,

\begin{array}{l} \frac{D_{ϕ_{s}} (P_{X} W_{Y | X} ∥ Q_{X} W_{Y | X})}{D_{ϕ_{s}} (P_{X} ∥ Q_{X})} \\ (281) & = \frac{χ^{2} (P_{X} W_{Y | X} ∥ (1 - s) P_{X} W_{Y | X} + s Q_{X} W_{Y | X})}{χ^{2} (P_{X} ∥ (1 - s) P_{X} + s Q_{X})} \\ (282) & \leq μ_{χ^{2}} ((1 - s) P_{X} + s Q_{X}, W_{Y | X}) \\ (283) & \leq μ_{χ^{2}} (W_{Y | X}), \end{array}

where (281) holds due to (18), and (283) is due to the definition in (105).

5.10. Proof of Proposition 3

The lower bound on the contraction coefficients in (108) and (109) is due to (276). The derivation of the upper bounds relies on [49] [Theorem 2.2], which states the following. Let

f : [0, \infty) \to R

be a three–times differentiable, convex function with

f (1) = 0

,

f^{″} (1) > 0

, and let the function

z : (0, \infty) \to R

defined as

z (t) : = \frac{f (t) - f (0)}{t}

, for all

t > 0

, be concave. Then,

\begin{matrix} μ_{f} (Q_{X}, W_{Y | X}) \leq \frac{f^{'} (1) + f (0)}{f^{″} (1) Q_{min}} \cdot μ_{χ^{2}} (Q_{X}, W_{Y | X}) . \end{matrix}

(284)

For

α \in (0, 1]

, let

z_{α, 1} : (0, \infty) \to R

and

z_{α, 2} : (0, \infty) \to R

be given by

\begin{matrix} z_{α, 1} (t) & : = \frac{k_{α} (t) - k_{α} (0)}{t}, t > 0, \end{matrix}

(285)

\begin{matrix} z_{α, 2} (t) & : = \frac{s_{α} (t) - s_{α} (0)}{t}, t > 0, \end{matrix}

(286)

with

k_{α}

and

s_{α}

in (97) and (98). Straightforward calculus shows that, for

α \in (0, 1]

and

t > 0

,

\begin{array}{l} (287) & \frac{1}{log e} z_{α, 1}^{″} (t) & = & - \frac{α^{2} + 2 α (1 - α) t}{t^{2} {[α + (1 - α) t]}^{2}} < 0, \\ \frac{1}{log e} z_{α, 2}^{″} (t) & = & - \frac{α^{2} [α + 2 (1 - α) t]}{t^{2} {[α + (1 - α) t]}^{2}} \\ (288) & - \frac{2 (1 - α)}{t^{3}} [{log}_{e} (1 + \frac{(1 - α) t}{α}) - \frac{(1 - α) t}{α + (1 - α) t} - \frac{{(1 - α)}^{2} t^{2}}{2 {[α + (1 - α) t]}^{2}}] . \end{array}

The first term on the right side of (288) is negative. For showing that the second term is also negative, we rely on the power series expansion

{log}_{e} (1 + u) = u - \frac{1}{2} u^{2} + \frac{1}{3} u^{3} - \dots

for

u \in (- 1, 1]

. Setting

u : = - \frac{x}{1 + x}

, for

x > 0

, and using Leibnitz theorem for alternating series yields

\begin{matrix} {log}_{e} (1 + x) = - {log}_{e} (1 - \frac{x}{1 + x}) > \frac{x}{1 + x} + \frac{x^{2}}{2 {(1 + x)}^{2}}, x > 0 . \end{matrix}

(289)

Consequently, setting

x : = \frac{(1 - α) t}{α} \in [0, \infty)

in (289), for

t > 0

and

α \in (0, 1]

, proves that the second term on the right side of (288) is negative. Hence,

z_{α, 1}^{″} (t), z_{α, 2}^{″} (t) < 0

, so both

z_{α, 1}, z_{α, 2} : (0, \infty) \to R

are concave functions.

In view of the satisfiability of the conditions of [49] [Theorem 2.2] for the f-divergences with

f = k_{α}

or

f = s_{α}

, the upper bounds in (108) and (109) follow from (284), and also since

\begin{array}{l} (290) & k_{α} (0) = 0, & k_{α}^{'} (1) = α log e, & k_{α}^{″} (1) = α^{2} log e, \\ (291) & s_{α} (0) = - (1 - α) log α, & s_{α}^{'} (1) = (2 α - 1) log e, & s_{α}^{″} (1) = (1 - 3 α + 3 α^{2}) log e . \end{array}

5.11. Proof of Proposition 4

In view of (24), we get

\begin{matrix} (292) & \frac{D (P_{Y} ∥ Q_{Y})}{D (P_{X} ∥ Q_{X})} & = \frac{\int_{0}^{1} χ^{2} (P_{Y} ∥ (1 - s) P_{Y} + s Q_{Y}) \frac{d s}{s}}{\int_{0}^{1} χ^{2} (P_{X} ∥ (1 - s) P_{X} + s Q_{X}) \frac{d s}{s}} \\ (293) & \leq \frac{\int_{0}^{1} μ_{χ^{2}} ((1 - s) P_{X} + s Q_{X}, W_{Y | X}) χ^{2} (P_{X} ∥ (1 - s) P_{X} + s Q_{X}) \frac{d s}{s}}{\int_{0}^{1} χ^{2} (P_{X} ∥ (1 - s) P_{X} + s Q_{X}) \frac{d s}{s}} \\ (294) & \leq sup_{s \in [0, 1]} μ_{χ^{2}} ((1 - s) P_{X} + s Q_{X}, W_{Y | X}) . \end{matrix}

In view of (119), the distributions of

X_{s}

and

Y_{s}

, and since

((1 - s) P_{X} + s Q_{X}) W_{Y | X} = (1 - s) P_{Y} + s Q_{Y}

holds for all

s \in [0, 1]

, it follows that

\begin{matrix} ρ_{m} (X_{s}; Y_{s}) = \sqrt{μ_{χ^{2}} ((1 - s) P_{X} + s Q_{X}, W_{Y | X})}, s \in [0, 1], \end{matrix}

(295)

which, from (292)–(295), implies that

\begin{matrix} sup_{s \in [0, 1]} ρ_{m} (X_{s}; Y_{s}) \geq \sqrt{\frac{D (P_{Y} ∥ Q_{Y})}{D (P_{X} ∥ Q_{X})}} . \end{matrix}

(296)

Switching

P_{X}

and

Q_{X}

in (292)–(294) and using the mapping

s \mapsto 1 - s

in (294) gives (due to the symmetry of the maximal correlation)

\begin{matrix} sup_{s \in [0, 1]} ρ_{m} (X_{s}; Y_{s}) \geq \sqrt{\frac{D (Q_{Y} ∥ P_{Y})}{D (Q_{X} ∥ P_{X})}}, \end{matrix}

(297)

and, finally, taking the maximal lower bound among those in (296) and (297) gives (120).

Author Contributions

Both coauthors contributed to this research work, and to the writing and proofreading of this article. The starting point of this work was in independent derivations of preliminary versions of Theorems 1 and 2 in two separate un-published works [24,25]. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Acknowledgments

Sergio Verdú is gratefully acknowledged for a careful reading, and well-appreciated feedback on the submitted version of this paper.

Conflicts of Interest

The authors declare no conflict of interest.

References

Kullback, S.; Leibler, R.A. On information and sufficiency. Ann. Math. Stat. 1951, 22, 79–86. [Google Scholar] [CrossRef]
Pearson, K. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. Lond. Edinb. Dublin Philos. Mag. J. Sci. 1900, 50, 157–175. [Google Scholar] [CrossRef] [Green Version]
Csiszár, I.; Shields, P.C. Information Theory and Statistics: A Tutorial. Found. Trends Commun. Inf. Theory 2004, 1, 417–528. [Google Scholar] [CrossRef] [Green Version]
Ali, S.M.; Silvey, S.D. A general class of coefficients of divergence of one distribution from another. J. R. Stat. Soc. 1966, 28, 131–142. [Google Scholar] [CrossRef]
Csiszár, I. Eine Informationstheoretische Ungleichung und ihre Anwendung auf den Bewis der Ergodizität von Markhoffschen Ketten. Publ. Math. Inst. Hungar. Acad. Sci. 1963, 8, 85–108. [Google Scholar]
Csiszár, I. Information-type measures of difference of probability distributions and indirect observations. Stud. Sci. Math. Hung. 1967, 2, 299–318. [Google Scholar]
Csiszár, I. On topological properties of f-divergences. Stud. Sci. Math. Hung. 1967, 2, 329–339. [Google Scholar]
Csiszár, I. A class of measures of informativity of observation channels. Period. Math. Hung. 1972, 2, 191–213. [Google Scholar] [CrossRef]
Van Erven, T.; Harremoës, P. Rényi divergence and Kullback–Leibler divergence. IEEE Trans. Inf. Theory 2014, 60, 3797–3820. [Google Scholar] [CrossRef] [Green Version]
Rényi, A. On measures of entropy and information. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics; University of California Press: Berkeley, CA, USA, 1961; pp. 547–561. [Google Scholar]
Liese, F.; Vajda, I. Convex Statistical Distances; Teubner-Texte Zur Mathematik: Leipzig, Germany, 1987. [Google Scholar]
Liese, F.; Vajda, I. On divergences and informations in statistics and information theory. IEEE Trans. Inf. Theory 2006, 52, 4394–4412. [Google Scholar] [CrossRef]
DeGroot, M.H. Uncertainty, information and sequential experiments. Ann. Math. Stat. 1962, 33, 404–419. [Google Scholar] [CrossRef]
Polyanskiy, Y.; Poor, H.V.; Verdú, S. Channel coding rate in the finite blocklength regime. IEEE Trans. Inf. Theory 2010, 56, 2307–2359. [Google Scholar] [CrossRef]
Sason, I.; Verdú, S. f-divergence inequalities. IEEE Trans. Inf. Theory 2016, 62, 5973–6006. [Google Scholar] [CrossRef]
Sason, I. On f-divergences: Integral representations, local behavior, and inequalities. Entropy 2018, 20, 383. [Google Scholar] [CrossRef] [Green Version]
Melbourne, J.; Madiman, M.; Salapaka, M.V. Relationships between certain f-divergences. In Proceedings of the 57th Annual Allerton Conference on Communication, Control and Computing, Urbana, IL, USA, 24–27 September 2019; pp. 1068–1073. [Google Scholar]
Melbourne, J.; Talukdar, S.; Bhaban, S.; Madiman, M.; Salapaka, M.V. The Differential Entropy of Mixtures: New Bounds and Applications. Available online: https://arxiv.org/pdf/1805.11257.pdf (accessed on 22 April 2020).
Audenaert, K.M.R. Quantum skew divergence. J. Math. Phys. 2014, 55, 112202. [Google Scholar] [CrossRef] [Green Version]
Gibbs, A.L.; Su, F.E. On choosing and bounding probability metrics. Int. Stat. Rev. 2002, 70, 419–435. [Google Scholar] [CrossRef] [Green Version]
Györfi, L.; Vajda, I. A class of modified Pearson and Neyman statistics. Stat. Decis. 2001, 19, 239–251. [Google Scholar] [CrossRef]
Le Cam, L. Asymptotic Methods in Statistical Decision Theory; Series in Statistics; Springer: New York, NY, USA, 1986. [Google Scholar]
Vincze, I. On the concept and measure of information contained in an observation. In Contributions to Probability; Gani, J., Rohatgi, V.K., Eds.; Academic Press: New York, NY, USA, 1981; pp. 207–214. [Google Scholar]
Nishiyama, T. A New Lower Bound for Kullback–Leibler Divergence Based on Hammersley-Chapman-Robbins Bound. Available online: https://arxiv.org/abs/1907.00288v3 (accessed on 2 November 2019).
Sason, I. On Csiszár’s f-divergences and informativities with applications. In Workshop on Channels, Statistics, Information, Secrecy and Randomness for the 80th birthday of I. Csiszár; The Rényi Institute of Mathematics, Hungarian Academy of Sciences: Budapest, Hungary, 2018. [Google Scholar]
Makur, A.; Polyanskiy, Y. Comparison of channels: Criteria for domination by a symmetric channel. IEEE Trans. Inf. Theory 2018, 64, 5704–5725. [Google Scholar] [CrossRef]
Simic, S. On a new moments inequality. Stat. Probab. Lett. 2008, 78, 2671–2678. [Google Scholar] [CrossRef]
Chapman, D.G.; Robbins, H. Minimum variance estimation without regularity assumptions. Ann. Math. Stat. 1951, 22, 581–586. [Google Scholar] [CrossRef]
Hammersley, J.M. On estimating restricted parameters. J. R. Stat. Soc. Ser. B 1950, 12, 192–240. [Google Scholar] [CrossRef]
Verdú, S. Information Theory, in preparation.
Wang, L.; Madiman, M. Beyond the entropy power inequality, via rearrangments. IEEE Trans. Inf. Theory 2014, 60, 5116–5137. [Google Scholar] [CrossRef] [Green Version]
Lewin, L. Polylogarithms and Associated Functions; Elsevier North Holland: Amsterdam, The Netherlands, 1981. [Google Scholar]
Marton, K. Bounding d¯-distance by informational divergence: A method to prove measure concentration. Ann. Probab. 1996, 24, 857–866. [Google Scholar] [CrossRef]
Marton, K. Distance-divergence inequalities. IEEE Inf. Theory Soc. Newsl. 2014, 64, 9–13. [Google Scholar]
Boucheron, S.; Lugosi, G.; Massart, P. Concentration Inequalities—A Nonasymptotic Theory of Independence; Oxford University Press: Oxford, UK, 2013. [Google Scholar]
Raginsky, M.; Sason, I. Concentration of Measure Inequalities in Information Theory, Communications and Coding: Third Edition. In Foundations and Trends in Communications and Information Theory; NOW Publishers: Boston, MA, USA; Delft, The Netherlands, 2018. [Google Scholar]
Csiszár, I. Sanov property, generalized I-projection and a conditional limit theorem. Ann. Probab. 1984, 12, 768–793. [Google Scholar] [CrossRef]
Clarke, B.S.; Barron, A.R. Information-theoretic asymptotics of Bayes methods. IEEE Trans. Inf. Theory 1990, 36, 453–471. [Google Scholar] [CrossRef] [Green Version]
Evans, R.J.; Boersma, J.; Blachman, N.M.; Jagers, A.A. The entropy of a Poisson distribution. SIAM Rev. 1988, 30, 314–317. [Google Scholar] [CrossRef] [Green Version]
Knessl, C. Integral representations and asymptotic expansions for Shannon and Rényi entropies. Appl. Math. Lett. 1998, 11, 69–74. [Google Scholar] [CrossRef] [Green Version]
Merhav, N.; Sason, I. An integral representation of the logarithmic function with applications in information theory. Entropy 2020, 22, 51. [Google Scholar] [CrossRef] [Green Version]
Cover, T.M.; Thomas, J.A. Elements of Information Theory, 2nd ed.; John Wiley & Sons: Hoboken, NJ, USA, 2006. [Google Scholar]
Csiszár, I. The method of types. IEEE Trans. Inf. Theory 1998, 44, 2505–2523. [Google Scholar] [CrossRef] [Green Version]
Corless, R.M.; Gonnet, G.H.; Hare, D.E.G.; Jeffrey, D.J.; Knuth, D.E. On the Lambert W function. Adv. Comput. Math. 1996, 5, 329–359. [Google Scholar] [CrossRef]
Tamm, U. Some refelections about the Lambert W function as inverse of x·log(x). In Proceedings of the 2014 IEEE Information Theory and Applications Workshop, San Diego, CA, USA, 9–14 February 2014. [Google Scholar]
Cohen, J.E.; Kemperman, J.H.B.; Zbăganu, G. Comparison of Stochastic Matrices with Applications in Information Theory, Statistics, Economics and Population Sciences; Birkhäuser: Boston, MA, USA, 1998. [Google Scholar]
Cohen, J.E.; Iwasa, Y.; Rautu, G.; Ruskai, M.B.; Seneta, E.; Zbăganu, G. Relative entropy under mappings by stochastic matrices. Linear Algebra Its Appl. 1993, 179, 211–235. [Google Scholar] [CrossRef] [Green Version]
Makur, A.; Zheng, L. Bounds between contraction coefficients. In Proceedings of the 53rd Annual Allerton Conference on Communication, Control and Computing, Urbana, IL, USA, 29 September–2 October 2015; pp. 1422–1429. [Google Scholar]
Makur, A. Information Contraction and Decomposition. Ph.D. Thesis, MIT, Cambridge, MA, USA, May 2019. [Google Scholar]
Polyanskiy, Y.; Wu, Y. Strong data processing inequalities for channels and Bayesian networks. In Convexity and Concentration; The IMA Volumes in Mathematics and its Applications; Carlen, E., Madiman, M., Werner, E.M., Eds.; Springer: New York, NY, USA, 2017; Volume 161, pp. 211–249. [Google Scholar]
Raginsky, M. Strong data processing inequalities and Φ-Sobolev inequalities for discrete channels. IEEE Trans. Inf. Theory 2016, 62, 3355–3389. [Google Scholar] [CrossRef] [Green Version]
Sason, I. On data-processing and majorization inequalities for f-divergences with applications. Entropy 2019, 21, 1022. [Google Scholar] [CrossRef] [Green Version]
Csiszár, I.; Körner, J. Information Theory: Coding Theorems for Discrete Memoryless Systems, 2nd ed.; Cambridge University Press: Cambridge, UK, 2011. [Google Scholar]
Burbea, J.; Rao, C.R. On the convexity of some divergence measures based on entropy functions. IEEE Trans. Inf. Theory 1982, 28, 489–495. [Google Scholar] [CrossRef]
Lin, J. Divergence measures based on the Shannon entropy. IEEE Trans. Inf. Theory 1991, 37, 145–151. [Google Scholar] [CrossRef] [Green Version]
Menéndez, M.L.; Pardo, J.A.; Pardo, L.; Pardo, M.C. The Jensen–Shannon divergence. J. Frankl. Inst. 1997, 334, 307–318. [Google Scholar] [CrossRef]
Topsøe, F. Some inequalities for information divergence and related measures of discrimination. IEEE Trans. Inf. Theory 2000, 46, 1602–1609. [Google Scholar] [CrossRef] [Green Version]
Nielsen, F. On a generalization of the Jensen–Shannon divergence and the Jensen–Shannon centroids. Entropy 2020, 22, 221. [Google Scholar] [CrossRef] [Green Version]
Asadi, M.; Ebrahimi, N.; Karazmi, O.; Soofi, E.S. Mixture models, Bayes Fisher information, and divergence measures. IEEE Trans. Inf. Theory 2019, 65, 2316–2321. [Google Scholar] [CrossRef]
Sarmanov, O.V. Maximum correlation coefficient (non-symmetric case). Sel. Transl. Math. Stat. Probab. 1962, 2, 207–210. (In Russian) [Google Scholar]
Gilardoni, G.L. Corrigendum to the note on the minimum f-divergence for given total variation. Comptes Rendus Math. 2010, 348, 299. [Google Scholar] [CrossRef]
Reid, M.D.; Williamson, R.C. Information, divergence and risk for binary experiments. J. Mach. Learn. Res. 2011, 12, 731–817. [Google Scholar]
Pardo, M.C.; Vajda, I. On asymptotic properties of information-theoretic divergences. IEEE Trans. Inf. Theory 2003, 49, 1860–1868. [Google Scholar] [CrossRef]

© 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).

Share and Cite

MDPI and ACS Style

Nishiyama, T.; Sason, I. On Relations Between the Relative Entropy and χ²-Divergence, Generalizations and Applications. Entropy 2020, 22, 563. https://doi.org/10.3390/e22050563

AMA Style

Nishiyama T, Sason I. On Relations Between the Relative Entropy and χ²-Divergence, Generalizations and Applications. Entropy. 2020; 22(5):563. https://doi.org/10.3390/e22050563

Chicago/Turabian Style

Nishiyama, Tomohiro, and Igal Sason. 2020. "On Relations Between the Relative Entropy and χ²-Divergence, Generalizations and Applications" Entropy 22, no. 5: 563. https://doi.org/10.3390/e22050563

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Menu

On Relations Between the Relative Entropy and χ²-Divergence, Generalizations and Applications

Abstract

1. Introduction

1.1. Paper Contributions

1.2. Paper Organization

2. Preliminaries and Notation

3. Relations between Divergences

3.1. Relations between the Relative Entropy and the Chi-Squared Divergence

3.2. Implications of Theorem 1

3.3. Monotonic Sequences of f-Divergences and an Extension of Theorem 1

3.4. On Probabilities and f-Divergences

4. Applications

4.1. Application of Corollary 3: Shannon Code for Universal Lossless Compression

4.2. Application of Theorem 2 in the Context of the Method of Types and Large Deviations Theory

4.3. Strong Data–Processing Inequalities and Maximal Correlation

5. Proofs

5.1. Proof of Theorem 1

5.2. Proof of Proposition 1

5.3. Proof of Theorem 2

5.4. Proof of Theorem 3

5.5. Proof of Theorem 4

5.6. Proof of Corollary 5

5.7. Proof of Theorem 5 and Corollary 6

5.8. Proof of Theorem 6

5.9. Proof of Corollary 7

5.10. Proof of Proposition 3

5.11. Proof of Proposition 4

Author Contributions

Funding

Acknowledgments

Conflicts of Interest

References

Share and Cite

Article Metrics

Article Access Statistics

Further Information

Guidelines

MDPI Initiatives

Follow MDPI