Publications · Journal

Approximation of probability density functions via location-scale finite mixtures in Lebesgue spaces

TrungTin Nguyen†, Faicel Chamroukhi, Hien D. Nguyen, Geoffrey J. McLachlan

† Corresponding author.

Commun. Stat. Theory Methods · Journal Communications in Statistics - Theory and Methods. Journal article (Vol. 52, Issue 14, pp. 5048-5059, 2022).

Abstract

The class of location-scale finite mixtures is of enduring interest both from applied and theoretical perspectives of probability and statistics. We establish and prove the following results: to an arbitrary degree of accuracy, (a) location-scale mixtures of a continuous probability density function (PDF) can approximate any continuous PDF, uniformly, on a compact set; and (b) for any finite p≥1, location-scale mixtures of an essentially bounded PDF can approximate any PDF in ℒp, in the ℒp norm.

1Normandie Univ, UNICAEN, CNRS, LMNO, 14000 Caen, France.
2School of Engineering and Mathematical Sciences. Department of Mathematics and Statistics, La Trobe University, Melbourne, Victoria, Australia.
3School of Mathematics and Physics, University of Queensland, St. Lucia, Brisbane, Australia.
∗∗Corresponding author.

Keywords: Mixture models, approximation theory, uniform approximation, probability density functions.

1 Introduction

Define (𝔼,∥⋅∥𝔼) to be a normed vector space (NVS), and let x∈(ℝn,∥⋅∥2), for some n∈ℕ, where ∥⋅∥2 is the Euclidean norm. Let f:ℝn→ℝ be a function satisfying f≥0 and ∫f​d​λ=1, where λ is the Lebesgue measure. We say that f is a probability density function (PDF) on the domain ℝn (which we will omit for brevity, from hereon in). Let g:ℝn→ℝ be another PDF and define the functional class ℳg=⋃m∈ℕℳmg, where

ℳmg={hmg:hmg​(⋅)=∑i=1mciσin​g​(⋅−μiσi),μi∈ℝn,σi∈ℝ+,c∈𝕊m−1,i∈[m]}​,

c⊤=(c1,…,cm), ℝ+=(0,∞),

𝕊m−1={c∈ℝm:∑i=1mci=1,ci≥0,i∈[m]}​,

[m]={1,…,m}, and (⋅)⊤ is the matrix transposition operator. We say that hmg∈ℳg is an m​-component location-scale finite mixture of the PDF g. The class ℳg has enjoyed enduring practical and theoretical interest throughout the years, as reported in the volumes of Everitt and Hand, (1981), McLachlan and Basford, (1988), Lindsay, (1995), McLachlan and Peel, (2000), Frühwirth-Schnatter, (2006), Mengersen et al., (2011), Frühwirth-Schnatter et al., (2019), and Nguyen et al., (2021).

We say that f is compactly supported on 𝕂⊂ℝn, if 𝕂 is compact and if 𝟏𝕂∁​f=0, where 𝟏𝕏 is the indicator function that takes value 1 when x∈𝕏, and 0 elsewhere, and where (⋅)∁ is the set complement operator (i.e. 𝕏∁=ℝn\𝕏). Here, 𝕏 is a generic subset of ℝn. Further, say that f∈ℒp​(𝕏) for any 1≤p<∞, if

‖f‖ℒp​(𝕏)=(∫|𝟏𝕏​f|p​d​λ)1/p<∞​,

and say that f∈ℒ∞​(𝕏), the class of essentially bounded measurable functions, if

‖f‖ℒ∞​(𝕏)=inf{a≥0:λ​({x∈𝕏:|f​(x)|>a})=0}<∞​,

where we call ∥⋅∥ℒp​(𝕏) the ℒp​-norm on 𝕏. Denote the class of all bounded functions on 𝕏 by

ℬ​(𝕏)={f∈ℒ∞​(𝕏):∃a∈[0,∞)​, such that ​|f​(x)|≤a,∀x∈𝕏}

and write

‖f‖ℬ​(𝕏)=supx∈𝕏|f​(x)|​.

For brevity, we shall write ℒp​(ℝn)=ℒp, ℬ​(ℝn)=ℬ, ‖f‖ℒp​(ℝn)=‖f‖ℒp, and ‖f‖ℬ​(ℝn)=‖f‖ℬ.

Lastly, we denote the class of continuous functions and uniformly continuous functions by 𝒞 and 𝒞u, respectively. The classes of bounded continuous function shall be denoted by 𝒞b=𝒞∩ℬ. Note that the class of continuous functions that vanish at infinity, defined as

𝒞0={f∈𝒞:∀ϵ>0,∃ a compact ​𝕂⊂ℝn​, such that ​‖f‖ℬ​(𝕂∁)<ϵ}​,

is a subset of 𝒞b.

An important characteristic of the class ℳg is its capability of approximating larger classes of PDFs in various ways. Motivated by the incomplete proofs of Xu et al., (1993, Lem 3.1) and Theorem 5 from Cheney and Light, (2000, Chapter 20), as well as the results of Nestoridis and Stefanopoulos, (2007), Bacharoglou, (2010), and Nestoridis et al., (2011), Nguyen et al., (2020) established and proved the following theorem regarding sequences of PDFs {hmg} from ℳg.

Theorem 1 (Theorem 5 from Nguyen et al., (2020)).

Let hmg∈ℳg denote an m​-component location finite mixture PDF. If we assume that f and g are PDFs and that g∈𝒞0, then the following statements are true.

  • (a)

    For any f∈𝒞0, there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞‖f−hmg‖ℒ∞=0​.
  • (b)

    For any f∈𝒞b, and compact set 𝕂⊂ℝn, there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞‖f−hmg‖ℒ∞​(𝕂)=0​.
  • (c)

    For any p∈(1,∞) and f∈ℒp, there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞‖f−hmg‖ℒp=0​.
  • (d)

    For any measurable f, there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞hmg=f​, almost everywhere.
  • (e)

    If ν is a σ​-finite Borel measure on ℝn, then for any ν​-measurable  f, there exists a sequence {hmg}​ℳg, such that

    limm→∞hmg=f​, almost everywhere, with respect to ​ν​.

Further, if we assume that

g∈{g∈𝒞0:∀x∈ℝn​, ​|g​(x)|≤θ1​(1+‖x‖2)−n−θ2​, ​(θ1,θ2)∈ℝ+2}​,

then the following is also true.

  • (f)

    For any f∈𝒞, there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞‖f−hmg‖ℒ1=0​.

The goal of this work is to seek the weakest set of assumptions in order to establish approximation theoretical results over the widest class of probability density problems, possible. In this paper, we establish Theorem 2 which improves upon Theorem 1 in a number of ways. More specifically, while statements (a), (c), (d), and (e) still hold under the same assumptions as in Theorem 1; statement (b) from Theorem 1 is improved to apply to a larger class of target function f∈𝒞, see more in statement (a) of Theorem 2; and statement (f) from Theorem 1 is drastically improved to apply to any f∈ℒ1 and g∈ℒ∞, see more in statement (b) of Theorem 2. We note in particular that our improvement with respect to statement (b) from Theorem 1 yields exactly the result of Theorem 5 from Cheney and Light, (2000, Chapter 20), which was incorrectly proved (see also DasGupta, (2008, Theorem 33.2)).

The remainder of the article progresses as follows. The main result of this paper is stated in Section 2. Technical preliminaries to the proof of the main result are presented in Section 3. The proof is then established in Section 4. Additional technical results required throughout the paper are reported in the Appendix A.

2 Main result

Theorem 2.

Let hmg∈ℳg denote an m​-component location finite mixture PDF. If we assume that f and g are PDFs, then the following statements are true.

  • (a)

    If f,g∈𝒞 and 𝕂⊂ℝn is a compact set, then there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞‖f−hmg‖ℬ​(𝕂)=0​.
  • (b)

    For p∈[1,∞), if f∈ℒp and g∈ℒ∞, then there exists a sequence {hmg}m=1∞⊂ℳg, such that

    limm→∞‖f−hmg‖ℒp=0​.

3 Technical preliminaries

Let f,g∈ℒ1, and denote the convolution of f and g by f⋆g=g⋆f. Further, we say that gk(⋅)=kng(k×⋅) (k∈ℝ+) is a dilate of g.

Notice that ℳmg can be parameterized via dilates. That is, we can write

ℳmg={hmg:hmg(⋅)=∑i=1mciking(ki×⋅−kiμi),μi∈ℝn,ki∈ℝ+,c∈𝕊m−1,i∈[m]},

where ki=1/σi.

Let 𝔽 be a subset of 𝔼, and denote the convex hull of 𝔽 by conv​(𝔽) is the smallest convex subset in 𝔼 that contains 𝔽 (cf. Brezis, 2010, Chapter 1). By definition, we may write

conv​(𝔽)={∑i∈[m]αi​fi:fi∈𝔽,α∈𝕊m−1,i∈[m],m∈ℕ}​,

where α⊤=(α1,…,αm).

Define the class of “basic” densities, which will serve as the approximation building blocks, as follows

𝒢g={kng(k×⋅−kμ),μ∈ℝn,k∈ℝ+},

and suppose that we can choose a suitable NVS (𝔼,∥⋅∥𝔼), such that 𝒢g⊂ℳg⊂𝔼. Then, by definition, it holds that ℳg is a convex hull of 𝒢g.

For u∈𝔼 and r>0, we define the open and closed balls of radius r, centered around u, by:

𝔹​(u,r)={v∈𝔼:‖u−v‖𝔼<r}​,

and

𝔹¯​(u,r)={v∈𝔼:‖u−v‖𝔼≤r}​,

respectively. For brevity, we also write 𝔹r=𝔹​(0,r) and 𝔹¯r=𝔹¯​(0,r). A set 𝔽⊂𝔼 is open, if for every u∈𝔽, there exists an r>0, such that 𝔹​(u,r)⊂𝔽. We say that 𝔽 is closed if its complement is open, and by definition, we say that 𝔼 and the empty set are both closed and open.

We call the smallest closed set containing 𝔽 its closure, and we denote it by 𝔽¯. A sequence {um}⊂𝔼 converges to u∈𝔼, if limm→∞‖um−u‖𝔼=0, and we denote it symbolically by limm→∞um=u. That is, for every ϵ>0, there exists an N​(ϵ)∈ℕ, such that m≥N​(ϵ) implies that ‖um−u‖𝔼<ϵ.

By Lemma 6, we can write the closure of 𝔽 as

𝔽¯={u∈𝔼:u=limm→∞um,um∈𝔽}

and hence

ℳg¯={h∈𝔼:h=limm→∞hmg,hmg∈ℳg}​.

Thus, by definition, it holds that ℳg¯ is a closed and convex subset of 𝔼.

If f∈𝒞 is a PDF on ℝn, we denote its support by

supp​f={x∈ℝn:f​(x)≠0}

and furthermore, we denote the set of compactly supported continuous functions by

𝒞c={f∈𝒞:supp​f​ is compact}​.

For open sets 𝕍⊂ℝn, we will write f≺𝕍 as shorthand for f∈𝒞c, 0≤f≤1, and supp​f⊂𝕍.

The following lemmas permit us to prove the primary technical mechanism that is used to prove our main result presented in Theorem 2.

Lemma 1.

Let f∈𝒞 be a PDF. Then, for every compact 𝕂⊂ℝn, we can choose h∈𝒞c, such that supp​h⊂𝔹r, 0≤h≤f, and h=f on 𝕂, for some r∈ℝ+.

Proof.

Since 𝕂 is bounded, there exists some r∈ℝ+, such that 𝕂⊂𝔹r. Lemma 10 implies that there exists a function u≺𝔹r, such that u​(x)=1, for all x∈𝕂. We can then set h=u​f to obtain the desired result of Lemma 1. ∎

Lemma 2.

Let h∈𝒞c, such that supp​h⊂𝔹r, 0≤h, and ∫h​d​λ≤1, and let g∈𝒞 be a PDF. Then, for any k∈ℝ+, there exists a sequence {hmg}m=1∞⊂ℳg, so that

limm→∞‖gk⋆h−hmg‖ℬ​(𝔹¯r)=0​. (1)

Furthermore, if g∈𝒞bu, we have the stronger result that

limm→∞‖gk⋆h−hmg‖ℬ=0​. (2)
Proof.

It suffices to show that given any r,k,ϵ∈ℝ+, there exists a sufficiently large m​(ϵ,r,k)∈ℕ such that for all m≥m​(ϵ,r,k), there exists a hmg∈ℳmg satisfying

‖gk⋆h−hmg‖ℬ​(𝔹¯r)<ϵ​. (3)

First, write

(gk⋆h)​(x) = ∫gk​(x−y)​h​(y)​d​λ​(y)=∫𝟏{y:y∈𝔹¯r}​gk​(x−y)​h​(y)​d​λ​(y)
= ∫𝟏{y:y∈𝔹¯r}​kn​g​(k​x−k​y)​h​(y)​d​λ​(y)=∫𝟏{z:z∈𝔹¯r​k}​g​(k​x−z)​h​(zy)​d​λ​(z)​,

where 𝔹¯r​k is a continuous image of a compact set, and hence is also compact (cf. Rudin, 1976, Theorem 4.14). By Lemma 11, for any δ>0, there exist κi∈ℝn (i∈[m−1], for some m∈ℕ), such that 𝔹¯r​k⊂⋃i=1m−1𝔹​(κi,δ/2). Further, if 𝔹iδ=𝔹r​kδ=𝔹¯r​k∩𝔹​(κi,δ/2), then 𝔹¯r​k=⋃i=1m−1𝔹iδ. We can hence obtain a disjoint covering of 𝔹¯r​k by taking 𝔸1δ=𝔹1δ, and 𝔸iδ=𝔹iδ\⋃j=1i−1𝔹jδ (i∈[m−1]) (cf. Cheney and Light, 2000, Chapter 24). Notice that 𝔹¯r​k=⋃i=1m−1𝔸iδ, each 𝔸iδ is a Borel set, and diam​(𝔸iδ)≤δ, by construction.

We shall denote the disjoint cover of 𝔹¯r​k by Πmδ={𝔸iδ}i=1m−1. We seek to show that there exists an m∈ℕ and Πmδ, such that

‖gk⋆h−∑i=1mci​kin​g​(ki​x−zi)‖ℬ​(𝔹¯r)<ϵ​,

where ki=k, ci=k−n​∫𝟏{z:z∈𝔸iδ}​h​(z/k)​d​λ​(z), and zi∈𝔸iδ, for i∈[m−1]. We then set zm=0 and cm=1−∑i=1m−1ci. Here, cm depends only on r and ϵ. Suppose that cm>0. Then, since g≠0, there exists some s∈ℝ+ such that Cs=supw∈𝔹¯sg​(w)>0. We can choose

km=min⁡{sr,(ϵ2​cm​Cs)1/n}​,

so that ∥g(km×⋅)∥ℬ​(𝔹¯r)≤Cs and

∥g(km×⋅)∥ℬ​(𝔹¯r)≤cm​ϵ​Cs2​cm​Cs=ϵ/2.

Moreover, if we assume that g∈𝒞bu, then there exists a constant C∈(0,∞) such that ‖g‖ℬ≤C. In this case, we can choose kmn=ϵ/(2​cm​C) to obtain

∥cmkmng(km×⋅−zm)∥ℬ≤ϵ/2.

Since 0≤h and ∫h​d​λ∈[0,1], the sum ∑i=1m−1ci satisfies the inequalities:

0≤∑i=1m−1ci = k−n​∑i=1m−1∫𝟏{z:z∈𝔸iδ}​h​(zk)​d​λ​(z)
= k−n​∫𝟏{z:z∈k​𝕂}​h​(zk)​d​λ​(z)=∫𝟏{x:x∈𝕂}​h​d​λ≤1​.

Thus, cm∈[0,1], and our construction of hmg implies that hmg=∑i=1mci​kin​g​(ki​x−zi)∈ℳmg.

We can then bound the left-hand side of (3) as follows:

‖gk⋆h−hmg‖ℬ​(𝔹¯r)
≤ ∥gk⋆h−∑i=1m−1ciking(ki×⋅−zi)∥ℬ​(𝔹¯r)+∥cmkmng(km×⋅−zm)∥ℬ​(𝔹¯r)
≤ ∥gk⋆h−∑i=1m−1ciking(ki×⋅−zi)∥ℬ​(𝔹¯r)+ϵ2
= ‖∫1{z:z∈𝔹¯r​k}​g​(k​x−z)​h​(zk)​d​λ​(z)−∑i=1m−1∫1{z:z∈𝔸iδ}​g​(k​x−z)​h​(zk)​d​λ​(z)‖ℬ​(𝔹¯r)
+ϵ2
≤ ∑i=1m−1∫1{z:z∈𝔸iδ}​|g​(k​x−z)−g​(k​x−zi)|​h​(zk)​d​λ​(z)+ϵ2​. (4)

Since x∈𝔹¯r, z∈𝔸iδ, and zi∈𝔹¯r​k, it holds that ‖k​x−zi‖2=‖k​x−z‖2≤2​r​k​, and

‖k​x−z−(k​x−zi)‖2=‖z−zi‖2≤diam​(𝔸iδ)≤δ​.

Note that g∈𝒞, and thus g is uniformly continuous on the compact set 𝔹¯2​r​k, implying that

|g​(k​x−z)−g​(k​x−zi)|≤w​(g,2​r​k,δ)​,

for each i∈[m−1], where

w(g,r,δ)=sup{|g(x)−g(y)|:∥x−y∥2≤δ and x,y∈𝔹¯r}

denotes a modulus of continuity. Since limδ→0w​(g,2​r​k,δ)=0 (cf. Makarov and Podkorytov, 2013, Theorem 4.7.3), we may choose a δ​(ϵ,r,k)>0, such that

w​(g,2​r​k,δ​(ϵ,r,k))<ϵ2​kn​.

We then proceed from (4) as follows:

‖gk⋆h−hmg‖ℬ​(𝔹¯r) ≤w​(g,2​r​k,δ​(ϵ,r,k))​∫𝟏{z:z∈𝔹¯r​k}​h​(zk)​d​λ​(z)+ϵ2
=w​(g,2​r​k,δ​(ϵ,r,k))​kn​∫h​d​λ+ϵ2
≤w​(g,2​r​k,δ​(ϵ,r,k))​kn+ϵ2<ϵ2+ϵ2=ϵ​. (5)

To conclude the proof of (1), it suffices to choose an appropriate sequence of partitions Πmδ​(ϵ,r,k), such that m≥m​(ϵ,r,k), for some sufficiently large m​(ϵ,r,k), so that (4) and (5) hold. This is possible via Lemma 11. When g∈𝒞bu, we notice that (4) and (5) both hold for all x∈ℝn. Thus, we have the stronger result of (2). ∎

We present the primary tools for proving Theorem (2) in the following pair of lemma. The first one in Lemma 3 permits the approximation of convolutions of the form gk⋆f in the ℒ1 functional space, and the second presented in Lemma 4 generalizes this first result to the spaces ℒp, where p∈[1,∞), under an essentially bounded assumption.

Lemma 3.

If f and g are PDFs in the NVS (ℒ1,∥⋅∥ℒ1), then ℳg⊂ℒ1 and gk⋆f∈ℒ1, for every k∈ℝ+. Furthermore, there exists a sequence {hmg}m=1∞⊂ℳg, such that

limm→∞‖gk⋆f−hmg‖ℒ1=0​.
Proof.

For any k∈ℝ+, we can show that gk∈ℒ1, since

‖gk‖ℒ1=∫gk​d​λ=∫kn​g​(k​x)​d​λ​(x)=∫g​d​λ=1​.

If hmg∈ℳmg, then hmg∈ℒ1, since it is a finite sum of functions in ℒ1, and thus, ℳg⊂ℒ1. Note that since f is a PDF, we have f∈ℒ1, and by Lemma 13, we also have that gk⋆f∈ℒ1. By Lemma 14, it then follows that

‖gk⋆f‖ℒ1 = ∫gk⋆f​d​λ
= ∫[∫gk​(x−y)​f​(y)​d​λ​(y)]​d​λ​(x)
= ∫[∫gk​(x−y)​d​λ​(x)]​f​(y)​d​λ​(y)
= ‖gk‖ℒ1​‖f‖ℒ1=1​

By definition of of the closure of ℳg in ℒ1, it suffices to show that for any k∈ℝ+, gk⋆f∈ℳg¯. We seek a contradiction by assuming that gk⋆f∉ℳg¯. Then, we can choose 𝔸=ℳg¯ and 𝔹={gk⋆f} so that 𝔸,𝔹⊂ℒ1 are nonempty convex subsets, such that 𝔸∩𝔹=∅. Furthermore, 𝔸 is closed and 𝔹 is compact. By Lemma 7, there exists a continuous linear functional ϕ∈ℒ1∗, such that ϕ​(v)<α<ϕ​(w), for all v∈𝔸 and w∈𝔹. By definition of 𝔹, for all v∈ℳg¯⊂ℒ1 we have

ϕ​(v)<α<ϕ​(gk⋆f)​.

By Lemma 9, with ϕ∈ℒ1∗, there exists a unique function u∈ℒ∞, such that, for all v∈ℒ1,

ϕ​(v)=∫u​(x)​v​(x)​d​λ​(x)​.

If we let v=gk(⋅−μ)∈ℳg¯⊂ℒ1, then we obtain the inequalities

supμ∈ℝn∫u​(x)​gk​(x−μ)​d​λ​(x)<α<∫u​(x)​(gk⋆f)​(x)​d​λ​(x)​.

The left-hand inequality can be reduced as follows:

α <∫u​(x)​(gk⋆f)​(x)​d​λ​(x)
=∫u​(x)​[∫gk​(x−μ)​f​(μ)​d​λ​(μ)]​d​λ​(x)
=∫f​(μ)​[∫u​(x)​gk​(x−μ)​d​λ​(x)]​d​λ​(μ)
<α​∫f​(μ)​d​λ​(μ)=α​,

where the third line is due to Lemma 14 and the final equality is because f is a PDF. This yields the sought contradiction. ∎

Lemma 4.

If f,g∈ℒ∞ are PDFs in the NVS (ℒ∞,∥⋅∥ℒp), for p∈[1,∞), then, ℳg⊂ℒp and gk⋆f∈ℒp, for any k∈ℝ+. Furthermore, there exists a sequence {hmg}m=1∞⊂ℳg, such that

limm→∞‖gk⋆f−hmg‖ℒp=0​.
Proof.

We obtain the result for p=1 via Lemma 3. Otherwise, since g∈ℒ1∩ℒ∞, we know that g∈ℒp and gk∈ℒp, for each k∈ℝ+, via Lemma 12. For any hmg∈ℳmg, we then have hmg∈ℒp via finite summation, and hence ℳg∈ℒp. Since f∈ℒ1, Lemma 13 implies that gk⋆f∈ℒp. By definition of the closure of ℳg, it suffices to show that gk⋆f∈ℳg¯, for any k∈ℝ+. This can be achieved by seeking a contradiction under the assumption that gk⋆f∉ℳg¯ and using Lemma 8 in the same manner as Lemma 9 is used in the proof of Lemma 3. ∎

4 Proof of main result

4.1 Proof of Theorem 2 (a)

To prove the statement (a) of Theorem 2, it suffices to show that there exists a sufficiently large m​(ϵ,𝕂)∈ℕ, such that for all m≥m​(ϵ,𝕂), there exists a hmg∈ℳmg, such that ‖f−hmg‖ℬ​(𝕂)<ϵ, for any ϵ>0 and compact set 𝕂⊂ℝn.

First, Lemma 1 implies that we can choose a h∈𝒞c, such that supp​h⊂𝔹¯r, 0≤h≤f, and h=f on 𝕂, for some r>0, where 𝕂⊂𝔹¯r. We then have ‖f−h‖ℬ​(𝕂)=0.

Since h∈𝒞c⊂𝒞bu, Lemma 5 and Corollary 1 then imply that there exists a k​(ϵ)∈ℝ+, such that for all k≥k​(ϵ), ‖h−gk⋆h‖ℬ​(𝕂)<ϵ/2. We shall assume that k≥k​(ϵ), from hereon in.

Lemma 2 then implies that there exists an m​(ϵ,r,k)∈ℕ, such that for any m≥m​(ϵ,r,k), there exists a hmg∈ℳmg, such that ‖gk⋆h−hmg‖ℬ​(𝕂)<‖gk⋆h−hmg‖ℬ​(𝔹¯r)<ϵ/2. The triangle inequality then completes the proof.

4.2 Proof of Theorem 2 (b)

To prove the statement (a) of Theorem 2, it suffices to show that there exists a sufficiently large m​(ϵ)∈ℕ, such that for all m≥m​(ϵ), there exists a hmg∈ℳmg, such that ‖f−hmg‖ℒp<ϵ, for any ϵ>0.

First, Lemma 5 and Corollary 1 imply that there exists a k​(ϵ)∈ℝ+, such that for any k≥k​(ϵ), it follows that ‖f−gk⋆f‖ℒp<ϵ/2. We shall assume k≥k​(ϵ), from hereon in.

Lemmas 3 and 4 imply that there exists an m​(ϵ)∈ℕ, such that for all m≥m​(ϵ), there exists a hmg∈ℳmg, such that ‖gk⋆f−hmg‖ℒp<ϵ/2. The triangle inequality then completes the proof.

Appendix A Technical results

We state a number of technical results that are used throughout the main text, in this Appendix. Sources for unproved results are provided at the end of the section.

Lemma 5.

Let {gk} be a sequence of PDFs in ℒ1, such that for every δ>0,

limk→∞∫𝟏{x:‖x‖2>δ}​gk​d​λ=0​.

Then, for f∈ℒp and p∈[1,∞),

limk→∞‖gk⋆f−f‖ℒp=0​.

Furthermore, for f∈𝒞b and compact 𝕂⊂ℝn,

limk→∞‖gk⋆f−f‖ℒ∞​(𝕂)=0​.

The sequences {gk} of Lemma 5 are often referred to as approximate identities or approximations of identity (cf. Makarov and Podkorytov, 2013, Sec. 7.6). A typical construction of approximate identities is to consider the sequence of dilations, of the form: gk(⋅)=kng(k×⋅), which permits the following corollary.

Corollary 1.

Let g be a PDF. Then, the sequence {gk:gk(⋅)=kng(k×⋅)} satisfies the hypothesis of Lemma 5 and hence permits its conclusion.

Lemma 6.

Let (𝔼,∥⋅∥𝔼) be an NVS, and let 𝔽⊂𝔼 and u∈𝔼. Then the following statements are equivalent: (a) u∈𝔽¯; (b) 𝔹​(u,r)∩𝔽≠∅, for all r>0; and (c) there exists a sequence {um}⊂𝔽 that converges to u.

Let 𝔼 be a locally convex linear topological space over ℝ and recall that a functional is a function defined on 𝔼 (or some subspace of 𝔼), with values in ℝ. We denote the due space of 𝔼 (the space of all continuous linear functions on 𝔼) by 𝔼∗.

Lemma 7 (Second geometric form of the Hahn-Banach theorem).

Let 𝔸,𝔹⊂𝔼 be two nonempty convex subsets, such that 𝔸∩𝔹≠∅. Assume that 𝔸 is closed and that 𝔹 is compact. Then, there exists a continuous linear functional ϕ∈𝔼∗, such that its corresponding hyperplane H={u∈𝔼:ϕ​(u)=α} (α∈ℝ) strictly separates 𝔸 and 𝔹. That is, there exists some ϵ>0, such that ϕ​(u)≤α−ϵ and ϕ​(v)≥α+ϵ, for all u∈𝔸 and v∈𝔹. Or, in other words, supu∈𝔸ϕ​(u)<infv∈𝔹ϕ​(v).

Lemma 8 (Riesz representation theorem for ℒp, p∈ℝ+).

If p∈ℝ+, and ϕ∈(ℒp)∗, then, there exists a unique function u∈ℒq, such that for all v∈ℒq,

ϕ​(v)=∫u​(x)​v​(x)​d​λ​(x)​,

where 1/p+1/q=1.

Lemma 9 (Riesz representation theorem for ℒ1).

If ϕ∈(ℒ1)∗, then there exists a unique u∈ℒ∞, such that for all v∈ℒ1,

ϕ​(v)=∫u​(x)​v​(x)​d​λ​(x)​.
Lemma 10.

Let 𝕍1,…,𝕍n be open subsets of ℝn, and let 𝕂 be a compact set, such that 𝕂⊂⋃i=1n𝕍i. Then, there exists functions hi≺𝕍i (i∈[n]), such that ∑i=1nhi​(x)=1, for all x∈𝕂. The set {hi} is referred to as the partition of unity on 𝕂, subordinated to the cover {𝕍i}.

Lemma 11.

If 𝕏⊂ℝn is bounded, then for any r>0, 𝕏 can be covered by ⋃i=1m𝔹​(xi,r), for some finite m∈ℕ, where xi∈ℝn and i∈[m].

Lemma 12.

If 1≤p≤q≤r≤∞, then ℒp∩ℒr⊂ℒq.

Lemma 13.

If f∈ℒp (1≤p≤∞) and g∈ℒ1, then f⋆g exists and we have ‖f⋆g‖ℒp≤‖f‖ℒp​‖f‖ℒ1. Furthermore, if p and q are such that 1/p+1/q=1, then f∈ℒp and g∈ℒq, then f⋆g exists, is bounded and uniformly continuous, and ‖f⋆g‖ℒ∞≤‖f‖ℒp​‖f‖ℒq. In particular, if p∈ℝ+, then f⋆g∈𝒞0.

Lemma 14 (Fubini’s Theorem).

Let (𝕏,𝒳,ν1) and (𝕐,𝒴,ν2) be σ​-finite measure spaces, and assume that f is a (𝒳×𝒴)​-measurable function on 𝕏×𝕐. If

∫𝕏[∫𝕐|f​(x,y)|​d​ν1​(x)]​d​ν2​(y)<∞​,

then

∫𝕏×𝕐|f|​d​(ν1×ν2) =∫𝕏[∫𝕐|f​(x,y)|​d​ν1​(x)]​d​ν2​(y)=∫𝕐[∫𝕏|f​(x,y)|​d​ν2​(y)]​d​ν1​(x)<∞​.

Sources for results

Lemma 5 appears in Makarov and Podkorytov, (2013, Thm. 9.3.3) and Cheney and Light, (2000, Ch. 20, Thm. 2). Corollary 1 is obtained from Cheney and Light, (2000, Ch. 20, Thm. 4). Lemmas 6, 12, and 13 are taken from Propositions 0.22, 6.10, and 8.8 Folland, (1999). Lemmas 7–9 appear in Brezis, (2010) as Theorems 1.7, 4.11, and 4.14, respectively. Lemmas 10 and 14 can be found in Rudin, (1987) as Theorems 2.13 and Theorem 8.8, respectively. Lemma 11 is obtained from Conway, (2012, Thm. 1.2.2).

Appendix B Acknowledgements

The authors would like to very much thank Pr. Eric Ricard for the interesting discussions with him and for his suggestions. TTN is supported by “Contrat doctoral” from the French Ministry of Higher Education and Research and by the French National Research Agency (ANR) grant SMILES ANR-18-CE40-0014. HDN and GJM are funded by Australian Research Council grant number DP180101192.

References

  • Bacharoglou, (2010) Bacharoglou, A. G. (2010). Approximation of probability distributions by convex mixtures of Gaussian measures. Proceeding of the American Mathematical Society, 138:2619–2628. ↩ 1
  • Brezis, (2010) Brezis, H. (2010). Functional Analysis, Sobolev Spaces and Partial Differential Equations. Springer, New York. ↩ 3 A
  • Cheney and Light, (2000) Cheney, W. and Light, W. (2000). A Course in Approximation Theory. Brooks/Cole, Pacific Grove. ↩ 1 3 A
  • Conway, (2012) Conway, J. B. (2012). A Course in Abstract Analysis. American Mathematical Society, Providence. ↩ A
  • DasGupta, (2008) DasGupta, A. (2008). Asymptotic Theory of Statistics and Probability. Springer, New York. ↩ 1
  • Everitt and Hand, (1981) Everitt, B. S. and Hand, D. J. (1981). Finite Mixture Distributions. Chapman and Hall, London. ↩ 1
  • Folland, (1999) Folland, G. B. (1999). Real Analysis: Modern Techniques and Their Applications. Wiley, New York. ↩ A
  • Frühwirth-Schnatter, (2006) Frühwirth-Schnatter, S. (2006). Finite Mixture and Markov Switching Models. Springer Science & Business Media. ↩ 1
  • Frühwirth-Schnatter et al., (2019) Frühwirth-Schnatter, S., Celeux, G., and Robert, C. P. (2019). Handbook of Mixture Analysis. CRC Press. ↩ 1
  • Lindsay, (1995) Lindsay, B. G. (1995). Mixture models: theory, geometry and applications. In NSF-CBMS Regional Conference Series in Probability and Statistics. ↩ 1
  • Makarov and Podkorytov, (2013) Makarov, B. and Podkorytov, A. (2013). Real Analysis: Measures, Integrals and Applications. Springer, London. ↩ 3 A
  • McLachlan and Basford, (1988) McLachlan, G. J. and Basford, K. E. (1988). Mixture Models: Inference and Applications to Clustering, volume 38. Marcel Dekker, New York. ↩ 1
  • McLachlan and Peel, (2000) McLachlan, G. J. and Peel, D. (2000). Finite Mixture Models. Wiley, New York. ↩ 1
  • Mengersen et al., (2011) Mengersen, K. L., Robert, C., and Titterington, M., editors (2011). Mixtures: Estimation and Applications. Wiley, Hoboken. ↩ 1
  • Nestoridis et al., (2011) Nestoridis, V., Schmutzhard, S., and Stefanopoulos, V. (2011). Universal series induced by approximate identities and some relevant applications. Journal of Approximation Theory, 163(12):1783–1797. ↩ 1
  • Nestoridis and Stefanopoulos, (2007) Nestoridis, V. and Stefanopoulos, V. (2007). Universal series and approximate identities. Technical report, University of Cyprus. ↩ 1
  • Nguyen et al., (2021) Nguyen, H. D., Nguyen, T., Chamroukhi, F., and McLachlan, G. J. (2021). Approximations of conditional probability density functions in Lebesgue spaces via mixture of experts models. Journal of Statistical Distributions and Applications, 8(1):13. ↩ 1
  • Nguyen et al., (2020) Nguyen, T. T., Nguyen, H. D., Chamroukhi, F., and McLachlan, G. J. (2020). Approximation by finite mixtures of continuous density functions that vanish at infinity. Cogent Mathematics & Statistics, 7(1):1750861. ↩ 1
  • Rudin, (1976) Rudin, W. (1976). Principles of Mathematical Analysis. McGraw-Hill, New York. ↩ 3
  • Rudin, (1987) Rudin, W. (1987). Real and complex analysis. McGraw-Hill, New York. ↩ A
  • Xu et al., (1993) Xu, Y., Light, W. A., and Cheney, E. W. (1993). Constructive methods of approximation by ridge functions and radial functions. Numerical Algoritms, 4:205–223. ↩ 1

Cite this paper

Please cite the published version. Venue: Communications in Statistics - Theory and Methods, Journal article (Vol. 52, Issue 14, pp. 5048-5059, 2022). DOI: 10.1080/03610926.2021.2002360. Official record: Taylor & Francis.

BibTeX
@article{nguyen2022approximation,
  title     = {Approximation of probability density functions via location-scale finite mixtures in Lebesgue spaces},
  author    = {Nguyen, TrungTin and Chamroukhi, Faicel and Nguyen, Hien D. and McLachlan, Geoffrey J.},
  journal   = {Communications in Statistics - Theory and Methods},
  volume    = {52}, number = {14}, pages = {5048--5059},
  year      = {2022}, publisher = {Taylor \& Francis},
  doi       = {10.1080/03610926.2021.2002360},
}