Publications · Conference

Approximation by Finite Mixtures of Continuous Density Functions That Vanish at Infinity

TrungTin Nguyen, Hien D. Nguyen, Faicel Chamroukhi, Geoffrey J. McLachlan

Cogent Math. & Stats. · Journal Cogent Mathematics & Statistics. Open-access journal article (Vol. 7, 2020).

Abstract

Given sufficiently many components, it is often cited that finite mixture models can approximate any other probability density function (pdf) to an arbitrary degree of accuracy. Unfortunately, the nature of this approximation result is often left unclear. We prove that finite mixture models constructed from pdfs in 𝒞0 can be used to conduct approximation of various classes of approximands in a number of different modes. That is, we prove approximands in 𝒞0 can be uniformly approximated, approximands in 𝒞b can be uniformly approximated on compact sets, and approximands in ℒp can be approximated with respect to the ℒp, for p∈[1,∞). Furthermore, we also prove that measurable functions can be approximated, almost everywhere.

Keywords— Approximation theory, probability density functions, finite mixture models, Riemann summation, uniform approximation.

1 Introduction

Let x be an element in the Euclidean space, defined by ℝn and the norm ∥⋅∥2, for some n∈ℕ. Let f:ℝn→ℝ be a function, such that f≥0, everywhere, and ∫f​d​λ=1, where λ is the Lesbegue measure. We say that f is a probability density function (pdf) on the domain ℝn (an expression that we will drop, from hereon in). Let g:ℝn→ℝ be another pdf, and for each m∈ℕ, define the functional class:

ℳmg = {h:h​(x)=∑i=1mci​1σin​g​(x−μiσi),μi∈ℝn​, ​σi∈ℝ+​, ​c∈𝕊m−1​, ​i∈[m]}​,

where c⊤=(c1,…,cm), ℝ+=(0,∞),

𝕊m−1={c∈ℝm:∑i=1mci=1​ and ​ci≥0,∀i∈[m]}​,

[m]={1,…,m}, and (⋅)⊤ is the matrix transposition operator. We say that any h∈ℳmg is a m​-component location-scale finite mixture of the pdf g.

The study of pdfs in the class ℳmg is an evergreen area of applied and technical research, in statistics. We point the interested reader to the many comprehensive books on the topic, such as [10],[35], [22], [20], [23], [14], [33], [24], and [15].

Much of the popularity of finite mixture models stem from the folk theorem, which states that for any density f, there exists an h∈ℳmg, for some sufficiently large number of components m∈ℕ, such that h approximates f arbitrarily closely, in some sense. Examples of this folk theorem come in statements such as: “provided the number of component densities is not bounded above, certain forms of mixture can be used to provide arbitrarily close approximation to a given probability distribution” [35, p. 50], “the [mixture] model forms can fit any distribution and significantly increase model fit” [37, p. 173], and “a mixture model can approximate almost any distribution” [39, p. 500]. Other statements conveying the same sentiment are reported in [28]. There is a sense of vagary in the reported statements, and little is ever made clear regarding the technical nature of the folk theorem.

In order to proceed, we require the following definitions. We say that f is compactly supported on 𝕂⊂ℝn, if 𝕂 is compact and if 𝟏𝕂∁​f=0, where 𝟏𝕏 is the indicator function that takes value 1 when x∈𝕏 and 0, elsewhere, and (⋅)∁ is the set complement operator (i.e., 𝕏∁=ℝn\𝕏). Here, 𝕏 is a generic subset of ℝn. Furthermore, we say that f∈ℒp​(𝕏) for any 1≤p<∞, if

‖f‖ℒp​(𝕏)=(∫|𝟏𝕏​f|p​d​λ)1/p<∞​,

and for p=∞, if

‖f‖ℒ∞​(𝕏)=inf{a≥0:λ​({x∈𝕏:|f​(x)|>a})=0}<∞,

where we call ∥⋅∥ℒp​(𝕏) the ℒp​-norm on 𝕏. When 𝕏=ℝn, we shall write ∥⋅∥ℒp​(ℝn)=∥⋅∥ℒp. In addition, we define the so-called Kullback-Leibler divergence, see [17], between any two pdfs f and g on 𝕏 as

KL𝕏​(f,g)=∫𝟏𝕏​f​log⁡(fg)​d​λ​.

In [28], the approximation of pdfs f by the class ℳmg was explored in a restrictive setting. Let {hmg} be a sequence of functions that draw elements from the nested sequence of sets {ℳmg} (i.e., h1g∈ℳ1g,h2g∈ℳ2g,…). The following result of [40] was presented in [28], along with a collection of its implications, such as the results of from [19] and [30].

Theorem 1 (Zeevi and Meir, 1997).

If

f∈{f:𝟏𝕂​f≥β​, ​β>0}∩ℒ2​(𝕂)

and g are pdfs and 𝕂 is compact, then there exists a sequence {hmg} such that

limm→∞‖f−hmg‖ℒ2​(𝕂)=0​ and ​limm→∞KL𝕂​(f,hmg)=0​.

Although powerful, this result is restrictive in the sense that it only permits approximation in the ℒ2 norm on compact sets 𝕂, and that the result only allows for approximation of functions f that are strictly positive on 𝕂. In general, other modes of approximation are desirable, in particular approximation in ℒp​-norm for p=1 or p=∞ are of interest, where the latter case is generally referred to as uniform approximation. Furthermore, the strict-positivity assumption, and the restriction on compact sets limits the scope of applicability of Theorem 1. An example of an interesting application of extensions beyond Theorem 1 is within the ℒ1​-norm approximation framework of [8].

Let g:ℝn→ℝ again be a pdf. Then, for each m∈ℕ, we define

𝒩mg ={h:h​(x)=∑i=1mci​1σin​g​(x−μiσi)​, ​μi∈ℝn​, ​σi∈ℝ+​, ​ci∈ℝ​, ​i∈[m]}​,

which we call the set of m​-component location-scale linear combinations of the pdf g. In the past, results regarding approximations of pdfs f via functions η∈𝒩mg have been more forthcoming. For example, in the case of g=ϕ, where

ϕ​(x)=(2​π)−n/2​exp⁡(−‖x‖22/2)​, (1)

is the standard normal pdf. Denoting the class of continuous functions with support on ℝn by 𝒞. We have the result that for every pdf f, compact set 𝕂⊂ℝn, and ϵ>0, there exists an m∈ℕ and h∈𝒩mϕ, such that ‖f−h‖ℒ∞​(𝕂)<ϵ [32, Lem. 1]. Furthermore, upon defining the set of continuous functions that vanish at infinity by

𝒞0 ={f∈𝒞:∀ϵ>0,∃ a compact ​𝕂⊂ℝn​, such that​‖f‖ℒ∞​(𝕂∁)<ϵ}​,

we also have the result: for every pdf f∈𝒞0 and ϵ>0, there exists an m∈ℕ and h∈𝒩mϕ, such that ‖f−h‖ℒ∞<ϵ [32, Thm. 2]. Both of the results from [32] are simple implications of the famous Stone-Weierstrass theorem (cf. [34] and [7]).

To the best of our knowledge, the strongest available claim that is made regarding the folk theorem, within a probabilistic or statistical context, is that of [6, Thm. 33.2]. Let {ηmg} be a sequence of functions that draw elements from the nested sequence of sets {𝒩mg}, in the same manner as {hmg}. We paraphrase the claim without loss of fidelity, as follows.

Claim 1.

If f,g∈𝒞 are pdfs and 𝕂⊂ℝn is compact, then there exists a sequence {ηmg}, such that

limm→∞‖f−ηmg‖ℒ∞​(𝕂)=0​.

Unfortunately, the proof of Claim 1 is not provided within [6]. The only reference of the result is to an undisclosed location in [4], which, upon investigation, can be inferred to be Theorem 5 of [4, Ch. 20]. It is further notable that there is no proof provided for the theorem. Instead, it is stated that the proof is similar to that of Theorem 1 in [4, Ch. 24], which is a reproduction of the proof for [38, Lem. 3.1].

There is a major problem in applying the proof technique of [38, Lem. 3.1] in order to prove Claim 1. The proof of [38, Lem. 3.1] critically depends upon the statement that “there is no loss of generality in assuming that f​(x)=0 for x∈ℝn\2​𝕂”. Here, for a∈ℝ+, a​𝕂={x∈ℝn:x=a​y​, ​y∈𝕂}. The assumption is necessary in order to write any convolution with f and an arbitrary continuous function as an integral over a compact domain, and then to use a Riemann sum to approximate such an integral. Subsequently, such a proof technique does not work outside the class of continuous functions that are compactly supported on a​𝕂. Thus, one cannot verify Claim 1 from the materials of [38], [4], and [6], alone.

Some recent results in the spirit of Claim 1 have been obtained by [27] and [26], using methods from the study of universal series (see for example in [25]).

Let

𝒲={f∈𝒞0:∑y∈ℤnsupx∈[0,1]n ​|f​(x+y)|<∞}

denote the so-called Wiener’s algebra (see, e.g., [11]) and let

𝒱 ={f∈𝒞0:∀x∈ℝn​, ​|f​(x)|≤β​(1+‖x‖2)−n−θ​, ​β,θ∈ℝ+}

be a class of functions with tails decaying at a faster rate than o​(‖x‖2n). In [26], it is noted that 𝒱⊂𝒲. Further, let

𝒞c={f∈𝒞:∃ a compact set ​𝕂​, such that ​𝟏𝕂∁​f=0}​,

denote the set of compactly supported continuous functions. The following theorem was proved in [27].

Theorem 2 (Nestoridis and Stefanopoulos, 2007, Thm. 3.2).

If g∈𝒱, then the following statements hold.

  • (a)

    For any f∈𝒞c, there exists a sequence {ηmg} (ηmg∈𝒩mg), such that

    limm→∞‖f−ηmg‖ℒ1+‖f−ηmg‖ℒ∞=0.
  • (b)

    For any f∈𝒞0, there exists a sequence {ηmg} (ηmg∈𝒩mg), such that

    limm→∞‖f−ηmg‖ℒ∞=0.
  • (c)

    For any 1≤p<∞ and f∈ℒp, there exists a sequence {ηmg} (ηmg∈𝒩mg), such that

    limm→∞‖f−ηmg‖ℒp=0.
  • (d)

    For any measurable f, there exists a sequence {ηmg} (ηmg∈𝒩mg), such that

    limm→∞ηmg=f​, almost everywhere.
  • (e)

    If ν is a σ​-finite Borel measure on ℝn, then for any ν​-measurable f, there exists a sequence {ηmg} (ηmg∈𝒩mg), such that

    limm→∞ηmg=f,

    almost everywhere, with respect to ν.

The result was then improved upon, in [26], whereupon the more general space 𝒲 was taken as a replacement for 𝒱, in Theorem 2. Denote the class of bounded continuous functions by 𝒞b=𝒞∩ℒ∞. The following theorem was proved in [26].

Theorem 3 (Nestoridis et al., 2011, Thm. 3.2).

If g∈𝒲, then the following statements are true.

  • (a)

    The conclusion of Theorem 2(a) holds, with 𝒞c replaced by 𝒞0∩ℒ1.

  • (b)

    The conclusions of Theorem 2(b)–(e) hold.

  • (c)

    For any f∈𝒞b and compact 𝕂⊂ℝn, there exists a sequence {ηmg}, such that

    limm→∞‖f−ηmg‖ℒ∞​(𝕂)=0.

Utilizing the techniques from [27], [1] proved a similar set of results to Theorem 2, under the restriction that f is a non-negative function with support ℝ, using g=ϕ (i.e. g has form (1), where n=1) and taking {hmϕ} as the approximating sequence, instead of {ηmg}. That is, the following result is obtained.

Theorem 4 (Bacharoglou, 2010, Cor. 2.5).

If f:ℝ→ℝ+∪{0}, then the following statements are true.

  • (a)

    For any pdf f∈𝒞c, there exists a sequence {hmϕ} (hmϕ∈ℳmϕ), such that

    limm→∞‖f−hmϕ‖ℒ1+‖f−hmϕ‖ℒ∞=0.
  • (b)

    For any f∈𝒞0, such that ‖f‖ℒ1≤1, there exists a sequence {hmϕ} (hmϕ∈ℳmϕ), such that

    limm→∞‖f−hmϕ‖ℒ∞=0.
  • (c)

    For any 1<p<∞ and f∈𝒞∩ℒp, such that ‖f‖ℒ1≤1, there exists a sequence {hmϕ} (hmϕ∈ℳmϕ), such that

    limm→∞‖f−hmϕ‖ℒp=0.
  • (d)

    For any measurable f, there exists a sequence {hmϕ} (hmϕ∈ℳmϕ), such that

    limm→∞hmϕ=f​, almost everywhere.
  • (e)

    For any pdf f∈𝒞, there exists a sequence {hmϕ} (hmϕ∈ℳmϕ), such that

    limm→∞‖f−hmϕ‖ℒ1=0.

To the best of our knowledge, Theorem 4 is the most complete characterization of the approximating capabilities of the mixture of normal distributions. However, it is restrictive in two ways. First, it does not permit characterization of approximation via the class ℳmg for any g except the normal pdf ϕ. Although ϕ is traditionally the most common choice for g in practice, the modern mixture model literature has seen the use of many more exotic component pdfs, such as the student-t pdf and its skew and modified variants (see, e.g., [29], [13], and [18]). Thus, its use is somewhat limited in the modern context. Furthermore, modern applications tend to call for n>1, further restricting the impact of the result as a theoretical bulwark for finite mixture modeling in practice. A remark in [1] states that the result can generalized to the case where g∈𝒱 instead of g=ϕ. However, no suggestions were proposed, regarding the generalization of Theorem 4 to the case of n>1.

In this article, we prove a novel set of results that largely generalize Theorem 4. Using techniques inspired by [9] and [4], we are able to obtain a set of results regarding the approximation capability of the class of m​-component mixture models ℳmg, when g∈𝒞0 or g∈𝒱, and for any n∈ℕ. By definition of 𝒱, the majority of our results extend beyond the proposed possible generalizations of Theorem 4.

The article proceeds as follows. Our main theorem is stated and its seperate parts are proved in the Section 2. Comments and discussion are provided in Section 3. Necessary technical lemmas and results are also included, for reference, in the Appendix.

2 Main result

The remainder of the article is devoted to proving the following theorem.

Theorem 5 (Main result).

If we assume that f and g are pdfs and that g∈𝒞0, then the following statements are true.

  • (a)

    For any f∈𝒞0, there exists a sequence {hmg} (hmg∈ℳmg), such that

    limm→∞‖f−hmg‖ℒ∞=0.
  • (b)

    For any f∈𝒞b and compact 𝕂⊂ℝn, there exists a sequence {hmg} (hmg∈ℳmg), such that

    limm→∞‖f−hmg‖ℒ∞​(𝕂)=0.
  • (c)

    For any 1<p<∞ and f∈ℒp, there exists a sequence {hmg} (hmg∈ℳmg), such that

    limm→∞‖f−hmg‖ℒp=0.
  • (d)

    For any measurable f, there exists a sequence {hmg} (hmg∈ℳmg), such that

    limm→∞hmg=f​, almost everywhere.
  • (e)

    If ν is a σ​-finite Borel measure on ℝn, then for any ν​-measurable f, there exists a sequence {hmg} (hmg∈ℳmg), such that

    limm→∞hmg=f,

    almost everywhere, with respect to ν.

If we assume instead that g∈𝒱, then the following statement is also true.

  • (f)

    For any f∈𝒞, there exists a sequence {hmg} (hmg∈ℳmg), such that

    limm→∞‖f−hmg‖ℒ1=0.

2.1 Technical preliminaries

Before we begin to prove the main theorem, we establish some technical results regarding our class of component densities 𝒞0. Let f,g∈ℒ1 and denote the convolution of f and g by f⋆g=g⋆f. Further, we denote the sequence of dilates of g by {gk:gk​(x)=kn​g​(k​x),k∈ℕ}. The following result is an alternative to Lemma 5 and Corollary 1. Here, we replace a boundedness assumption on the approximand, in the aforementioned theorem by a vanishing at infinity assumption, instead.

Lemma 1.

Let g be a pdf and f∈𝒞0, such that ‖f‖ℒ∞>0. Then,

limk→∞‖gk⋆f−f‖ℒ∞=0​.
Proof.

It suffices to show that for any ϵ>0, there exists a k​(ϵ)∈ℕ, such that ‖gk⋆f−f‖ℒ∞<ϵ, for all k≥k​(ϵ). By Lemma 6, f∈𝒞b, and thus ‖f‖ℒ∞<∞. By making the substitution z=k​x, we obtain for each k

∫gk​(x)​d​λ=∫kn​g​(k​x)​d​λ=∫g​(z)​d​λ=1​.

By Corollary 1, we obtain limk→∞∫𝟏{x:‖x‖2>δ}​gk​d​λ=0 and thus we can choose a k​(ϵ), such that

∫𝟏{x:‖x‖2>δ}​gk​d​λ<ϵ4​‖f‖ℒ∞​.

Since g is a pdf, we have

|(gk⋆f)​(x)−f​(x)| =|∫gk​(y)​[f​(x−y)−f​(x)]​d​λ​(y)|
≤∫gk​(y)​|f​(x−y)−f​(x)|​d​λ​(y)​.

By uniform continuity, for any ϵ>0, there exists a δ​(ϵ)>0 such that |f​(x−y)−f​(x)|<ϵ/2, for any x,y∈ℝn, such that ‖y‖2<δ​(ϵ) (Lemma 6). Thus, on the one hand, for any δ​(ϵ), we can pick a k​(ϵ) such that

∫𝟏{y:‖y‖2>δ​(ϵ)}​gk​(y)​|f​(x−y)−f​(x)|​d​λ​(y)
≤2​‖f‖ℒ∞​∫𝟏{y:‖y‖2>δ​(ϵ)}​gk​d​λ
≤2​‖f‖ℒ∞×ϵ4​‖f‖ℒ∞=ϵ2​, (2)

and on the other hand

∫𝟏{y:‖y‖2≤δ​(ϵ)}​gk​(y)​|f​(x−y)−f​(x)|​d​λ​(y)
≤ϵ2​∫𝟏{y:‖y‖2≤δ​(ϵ)}​gk​d​λ
≤ϵ2×1=ϵ2​. (3)

The proof is completed by summing (2) and (3). ∎

Lemma 2.

If f∈𝒞0 is such that f≥0, and ϵ>0, then there exists a h∈𝒞c, such that 0≤h≤f, and

‖f−h‖ℒ∞<ϵ
Proof.

Since f∈𝒞0, there exists a compact 𝕂⊂ℝn such that ‖f‖ℒ∞​(𝕂∁)<ϵ/2. By Lemma 7, there exists some g∈𝒞c, such that 0≤g≤1 and 𝟏𝕂​g=1. Let h=g​f, which implies that h≥0 and 0≤h≤f. Furthermore, notice that 𝟏𝕂​(f−h)=0 and ‖h‖ℒ∞≤‖f‖ℒ∞, by construction. The proof is completed by observing that

‖f−h‖ℒ∞ =‖f−h‖ℒ∞​(𝕂∁)
≤‖f‖ℒ∞​(𝕂∁)+‖h‖ℒ∞​(𝕂∁)
≤2​‖f‖ℒ∞​(𝕂∁)<ϵ​.

∎

For any δ>0, uniformly continuous function f, let

w​(f,δ)=sup{x,y∈ℝn:‖x−y‖2≤δ}|f​(x)−f​(y)|

denote the modulus of continuity of f. Furthermore, define the diameter of a set 𝕏⊂ℝn by diam​(𝕏)=supx,y∈𝕏‖x−y‖2 and denote an open ball, centered at x∈ℝn with radius r>0 by 𝔹​(x,r)={y∈ℝn:‖x−y‖2<r}.

Notice that the class ℳmg can be parameterized as

ℳmg ={h:h(x)=∑i=1mciking(kix−zi),
 zi∈ℝn, ki∈ℝ+, c∈𝕊m−1, i∈[m]},

where ki=1/σi and zi=μi/σi. The following result is the primary mechanism that permits us to construct finite mixture approximations for convolutions of form gk⋆f. The argument motivated by the approaches taken in Theorem 1 in [4, Ch. 24], [27, Lem. 3.1], and [26, Thm. 3.1].

Lemma 3.

Let f∈𝒞 and g∈𝒞0 be pdfs. Furthermore, let 𝕂⊂ℝn be compact and h∈𝒞c, where 𝟏𝕂∁​h=0 and 0≤h≤f. Then for any k∈ℕ, there exists a sequence {hmg}, such that

limm→∞‖gk⋆h−hmg‖ℒ∞=0.
Proof.

It suffices to show that for any k∈ℕ and ϵ>0, there exists a sufficiently large enough m​(ϵ)∈ℕ so that for all m≥m​(ϵ),hmg∈ℳmg such that

‖gk⋆h−hmg‖ℒ∞<ϵ. (4)

For any k∈ℕ, we can write

(gk⋆h)​(x) =∫gk​(x−y)​h​(y)​d​λ​(y)
=∫𝟏{y:y∈𝕂}​gk​(x−y)​h​(y)​d​λ​(y)
=∫𝟏{y:y∈𝕂}​kn​g​(k​x−k​y)​h​(y)​d​λ​(y)
=∫𝟏{z:z∈k​𝕂}​g​(k​x−z)​h​(zk)​d​λ​(z)​.

Here, k​𝕂 is continuous image of a compact set, and hence is compact (cf. [31, Thm. 4.14]). By Lemma 8, for any δ>0, there exists κi∈ℝn (i∈[m−1], m∈ℕ), such that k​𝕂⊂⋃i=1m−1𝔹​(κi,δ/2). Further, if 𝔹iδ=k​𝕂∩𝔹​(κi,δ/2), then we have k​𝕂=⋃i=1m−1𝔹iδ. We can obtain a disjoint covering of k​𝕂 by taking 𝔸1δ=𝔹1 and 𝔸iδ=𝔹iδ\⋃j=1i−1𝔹jδ (i∈[m−1]) and noting that k​𝕂=⋃i=1m−1𝔸iδ, by construction (cf. [4, Ch. 24]). Furthermore, each 𝔸iδ is a Borel set and diam​(𝔸iδ)≤δ.

For convenience, let Πmδ={𝔸iδ:i∈[m−1]} denote the disjoint covering, or partition, of k​𝕂. We seek to show that there exists an m∈ℕ and Πmδ, such that

‖gk⋆h−∑i=1mci​kin​g​(ki​x−zi)‖ℒ∞<ϵ​,

where ki=k,

ci=k−n​∫𝟏{z:z∈𝔸iδ}​h​(z/k)​d​λ​(z),

and zi∈𝔸iδ, for i∈[m−1].

Further, zm∈𝔸m−1δ and cm=1−∑i=1m−1ci, with km chosen as follows. By Lemma 6, g≤C<∞ for some positive C. Then, ‖cm​kmn​g​(km​x−zm)‖ℒ∞≤cm​kmn​C. We may choose km so that kmn=ϵ/(2​cm​C), so that

‖cm​kmn​g​(km​x−zm)‖ℒ∞≤ϵ2​.

Since 0≤h≤f, the sum of ci (i∈[m−1]) satisfies the inequality

∑i=1m−1ci =k−n​∑i=1m−1∫𝟏{z:z∈𝔸iδ}​h​(zk)​d​λ
=k−n​∫𝟏{z:z∈k​𝕂}​h​(zk)​d​λ
=∫𝟏{x:x∈𝕂}​h​d​λ≤∫𝟏{x:x∈𝕂}​f​d​λ≤∫f​d​λ=1​.

Thus, 0≤cm≤1, and our construction implies that hmg∈ℳmg​, where

hmg​(x)=∑i=1mci​kin​g​(ki​x−zi)​∀x∈ℝn​.

We can bound the left-hand side of (4) as follows:

‖gk⋆h−hgm‖ℒ∞
≤‖(gk⋆h)​(x)−∑i=1m−1ci​kin​g​(ki​x−zi)‖ℒ∞
 +‖cm​kmn​g​(km​x−zm)‖ℒ∞
≤‖(gk⋆h)​(x)−∑i=1m−1ci​kin​g​(ki​x−zi)‖ℒ∞+ϵ2
=∥∫𝟏{z:z∈k​𝕂}g(kx−z)h(zk)dλ(z)
 −∑i=1m−1∫𝟏{z:z∈𝔸iδ}​g​(k​x−zi)​h​(zk)​d​λ​(z)∥ℒ∞+ϵ2
≤∑i=1m−1∫𝟏{z:z∈𝔸iδ}​‖g​(k​x−z)−g​(k​x−zi)‖ℒ∞​h​(zk)​d​λ​(z)+ϵ2​. (5)

Since

‖k​x−z−(k​x−zi)‖2=‖z−zi‖2≤diam​(𝔸iδ)≤δ​,

we have |g​(k​x−z)−g​(k​x−zi)|≤w​(g,δ), for each i∈[m−1]. Since limδ→0w​(g,δ)=0 (cf. [21, Thm. 4.7.3]), we may choose a δ​(ϵ)>0 so that w​(g,δ​(ϵ))<ϵ/(2​kn). We may proceed from (2.1) as follows:

‖gk⋆h−hgm‖ℒ∞ ≤w​(g,δ​(ϵ))​∫𝟏{z:z∈k​𝕂}​h​(zk)​d​λ+ϵ2
=w​(g,δ​(ϵ))​kn​∫h​d​λ+ϵ2
≤w​(g,δ​(ϵ))​kn+ϵ2
<ϵ2+ϵ2=ϵ​. (6)

To conclude the proof, it suffices to choose an appropriate sequence of partitions Πmδ​(ϵ),m≥m​(ϵ), for some large but finite m​(ϵ), so that (2.1) and (6) hold, which is possible by Lemma 8. ∎

For any r∈ℕ, let 𝔹¯r={x∈ℝn:‖x‖2≤r} be a closed ball of radius r, centered at the origin.

Lemma 4.

If f∈ℒ1, such that f≥0, then

limr→∞‖f−𝟏𝔹¯r​f‖ℒ1=0​.
Proof.

By construction, each element of the sequence {𝟏𝔹¯r​f} (r∈ℕ) is measurable, 0≤𝟏𝔹¯r​f≤f, and

limr→∞𝟏𝔹¯r​f=f​,

point-wise. We obtain our conclusion via the Lesbegue dominated convergence theorem. ∎

2.2 Proof of Theorem 5(a)

We now proceed to prove each of the parts of Theorem 5. To prove Theorem 5(a) it suffices to show that for every ϵ>0, there exists a hmg∈ℳmg, such that ‖f−hmg‖ℒ∞<ϵ​.

Start by applying Lemma 2 to obtain h∈𝒞c, such that 0≤h≤f and ‖f−h‖ℒ∞<ϵ/2. Then, we have

‖f−hmg‖ℒ∞ ≤‖f−h‖ℒ∞+‖h−hmg‖ℒ∞
<ϵ2+∥​h−hmg∥ℒ∞​. (7)

The goal is to find a hmg, such that ‖h−hmg‖ℒ∞<ϵ/2. Since h∈𝒞c, we may find a compact 𝕂⊂ℝn such that ‖h‖ℒ∞​(𝕂∁)=0. Apply Lemma 1 to show the existence of a k​(ϵ), such that

‖h−gk⋆h‖ℒ∞<ϵ4​,

for all k≥k​(ϵ). With a fixed k=k​(ϵ), apply Lemma 3 to show that there exists a hmg∈ℳmg, such that

‖gk​(ϵ)⋆h−hmg‖ℒ∞<ϵ4​.

By the triangle inequality, we have

‖h−hmg‖ℒ∞ ≤‖h−gk​(ϵ)⋆h‖ℒ∞+‖gk​(ϵ)⋆h−hmg‖ℒ∞
<ϵ4+ϵ4=ϵ2​. (8)

The proof is complete by substitution of (8) into (7).

2.3 Proof of Theorem 5(b)

For any ϵ>0 and compact 𝕂⊂ℝn, it suffices to show that there exists a sufficiently large enough m​(ϵ)∈ℕ so that for all m≥m​(ϵ),hmg∈ℳmg, such that ‖f−hmg‖ℒ∞​(𝕂)<ϵ.

By Lemma 5, we can find a k​(ϵ,𝕂)∈ℕ, such that

‖f−gk⋆f‖ℒ∞​(𝕂)<ϵ3​, (9)

for every k≥k​(ϵ,𝕂). Since g∈𝒞0, ‖g‖ℒ∞≤C<∞ for some positive C, by Lemma 6. For any k,r∈ℕ, via Young’s convolution inequality:

‖gk⋆f−gk⋆(𝟏𝔹¯r​f)‖ℒ∞≤kn​C​∫(𝟏𝔹¯r∁​f)​d​λ=kn​C​‖f−𝟏𝔹¯r​f‖ℒ1​. (10)

For fixed k, we may choose r​(ϵ,𝕂)∈ℕ, using Lemma 4, so that ‖f−𝟏𝔹¯r​f‖ℒ1≤ϵ/(3​kn​C) and thus the final term of (10) is bounded from above by ϵ/3 for all r≥r​(ϵ,𝕂). Thus, for k=k​(ϵ,𝕂) and, r≥r​(ϵ,𝕂)

‖gk​(ϵ,𝕂)⋆f−gk​(ϵ,𝕂)⋆(𝟏𝔹¯r​(ϵ,𝕂)​f)‖ℒ∞≤ϵ3​. (11)

Using Lemma 3, with approximand 𝟏𝔹¯r​(ϵ,𝕂)​f, component density g, compact set 𝔹¯r​(ϵ,𝕂), h=𝟏𝔹¯r​(ϵ,𝕂)​f, and with k=k​(ϵ,𝕂) fixed, we have the existence of a density hmg∈ℳmg,m≥m​(ϵ)∈ℕ, such that

‖gk​(ϵ,𝕂)⋆(𝟏𝔹¯r​(ϵ,𝕂)​f)−hmg‖ℒ∞≤ϵ3​. (12)

We obtain the desired result by combining (9), (11), and (12), via the triangle inequality.

2.4 Proof of Theorem 5(c)

The technique used to prove Theorem 5(c) is different to those used in the previous sections. Here, we use a result of [9] that generalizes the classic Barron-Jones Hilbert space approximation result (cf. [16] and [2]) to Banach spaces.

To prove Theorem 5(c), it suffices to show that for every ϵ>0, there exists a sufficiently large enough m​(ϵ)∈ℕ so that for all m≥m​(ϵ),hmg∈ℳmg such that ‖f−hmg‖ℒp<ϵ. Begin by applying Corollary 1 to obtain a k​(ϵ), such that

‖f−gk⋆f‖ℒp<ϵ2 (13)

for all k≥k​(ϵ).

For some pdf g and fixed k∈ℕ, let us define the class

𝒢gk={h:h​(x)=kn​g​(k​x−k​μ)​, ​μ∈ℝn}​,

write the m​-point convex hull of 𝒢gk as

Convm​(𝒢gk)={h:h=∑i=1mci​gi​, ​gi∈𝒢gk​, ​c∈𝕊m−1​, ​i∈[m]}​,

and call Conv∞​(𝒢gk)=Conv​(𝒢gk) the convex hull of 𝒢gk. We further say that Conv¯​(𝒢gk) is the closure of Conv​(𝒢gk).

Because g is a pdf, g∈𝒞0⊂𝒞b, and 𝒞b⊂ℒ∞, we observe that g∈ℒ1∩ℒ∞. Thus, g∈ℒp, for any 1<p<∞, by Lemma 9. Since g is a pdf and f∈ℒp, we have the existence of gk⋆f and the fact that ‖gk⋆f‖ℒp is finite.

Furthermore, for any ψ∈𝒢gk, since g∈ℒp and by definition of 𝒢gk, we have ‖ψ‖ℒp≤kn/p​‖g‖ℒp. Thus, we have

‖ψ−gk⋆f‖ℒp≤‖ψ‖ℒp+‖gk⋆f‖ℒp≤K​, (14)

by choosing K=kn/p​‖g‖ℒp+‖gk⋆f‖ℒp>0.

Following [36], we can write the closure of 𝒢gk as

Conv¯​(Ggk)={h:h​(x)=∫kn​g​(k​x−k​μ)​f​(μ)​d​λ​(μ),f​ is a pdf}​,

and thus we immediately have gk⋆f∈Conv¯​(Ggk). Combined with (14), we can apply Lemma 11 to obtain the conclusion that there exists a function hmg∈Convm​(𝒢gk​(ϵ))⊂ℳmg, such that

‖hmg−gk​(ϵ)⋆f‖ℒp≤K​Cpm1−1/α​,

where α=min⁡{p,2} and Cp is a finite constant. Since p>1, m1−1/α is strictly increasing, and hence we can choose an m​(ϵ)∈ℕ, such that for all m≥m​(ϵ),

‖hmg−gk​(ϵ)⋆f‖ℒp≤ϵ2​. (15)

The proof is then completed by combining (13) and (15) via the triangle inequality.

2.5 Proof of Theorem 5(d) and Theorem 5(e)

By Theorem 5(a), there exists a sequence {hmg} that uniformly converges to f, as m→∞. Thus, by Lemma 12, {hmg} almost uniformly converges to f and also converges almost everywhere, to f, with respect to any measure ν. We prove Theorem 5(d) by setting ν=λ, and we prove Theorem 5(e) by not specifying ν.

2.6 Proof of Theorem 5(f)

It suffices to show that for any ϵ>0, there exists a sufficiently large enough m​(ϵ)∈ℕ so that for all m≥m​(ϵ),hmg∈ℳmg, where g∈𝒱, such that ‖f−hmg‖ℒ1<ϵ. Begin by applying Lemma 4 in order to find a r​(ϵ)∈ℕ, for any ϵ>0, such that for all r≥r​(ϵ),

‖f−𝟏𝔹¯r​f‖ℒ1≤ϵ24<ϵ2​, (16)

where 0≤𝟏𝔹¯r​f≤f, and 𝟏𝔹¯r​f∈𝒞c with compact support 𝔹¯r.

Let 𝕂=𝔹¯r and apply the triangle inequality to obtain

‖f−hmg‖ℒ1 ≤‖f−𝟏𝕂​f‖ℒ1+‖𝟏𝕂​f−hmg‖ℒ1
≤ϵ2+‖𝟏𝕂​f−hmg‖ℒ1​.

Hence we need to show that there exists a function hmg∈ℳmg, such that

‖𝟏𝕂​f−hmg‖ℒ1≤ϵ2​.

Since g∈𝒱 and gk​(x)=kn​g​(k​x), by substitution, we have

gk​(x)≤β​k−θ(k−1+‖x‖2)n+θ​, (17)

where β,θ>0 are independent of k. By Lemma 5 and Corollary 1, we can obtain a k1​(ϵ), such that for all k≥k1​(ϵ),

‖𝟏𝕂​f−gk⋆(𝟏𝕂​f)‖ℒ1≤ϵ4​. (18)

Suppose that γ>1 and let

𝕂k={x∈ℝn:dist​(x,𝕂)≤k−γ},

where

dist​(x,𝕏)=inf{‖x−y‖2:y∈𝕏}​.

By construction, λ​(𝕂k)=λ​(𝕂)+O​(k−γ) and thus there exists a k2 such that λ​(𝕂k)≤λ​(𝕂)+1, for any k≥k2.

For any k>k2, we can show that

‖gk⋆(𝟏𝕂​f)−hm−1g‖ℒ1​(𝕂k)<ϵ8​. (19)

To do so, firstly, for any x∈ℝn,

gk⋆(𝟏𝕂​f) =∫𝟏𝕂​gk​(x−y)​f​(y)​d​λ​(y)
=∫𝟏k​𝕂​g​(k​x−z)​f​(zk)​d​λ​(z)​.

To obtain a Riemann sum approximation of gk⋆(𝟏𝕂​f), we use an argument analogous to that of Lemma 3. That is, we partition k​𝕂 into m−1 disjoint Borel sets Πm={𝔸1,…,𝔸m−1}, and we approximate gk⋆(𝟏𝕂​f) by a hm−1g∈ℳm−1g, where for each i∈[m−1], ki=k, zi∈𝔸i, and

ci=k−n​∫𝟏𝔸i​f​(zk)​d​λ​(z)​.

Define km∈ℝ+, zm∈ℝn, and cm=1−∑i=1m−1ci, where

cm=∫f​d​λ−∫𝟏𝕂​f​d​λ=‖f−𝟏𝕂​f‖ℒ1≤ϵ24 (20)

by (16). Then, by a similar argument to Lemma 3, ci≥0 for all i∈[m] and ∑i=1mci=1. Thus, we may define an element hmg∈ℳmg via the parameters above.

For sufficiently large k≥k2, we use Lemma 3 to show that

‖gk⋆(𝟏𝕂​f)−hm−1g‖ℒ∞​(𝕂k)<ϵ8​(λ​(𝕂)+1)​,

which implies

‖gk⋆(𝟏𝕂​f)−hm−1g‖ℒ1​(𝕂k) <∫𝟏𝕂k​ϵ8​(λ​(𝕂)+1)​d​λ
<ϵ​λ​(𝕂k)8​(λ​(𝕂)+1)<ϵ8​, (21)

and thus (19) is proved. Using (19), we write

‖gk⋆(𝟏𝕂​f)−hmg‖ℒ1
= ‖gk⋆(𝟏𝕂​f)−hm−1g−cm​kmn​g​(km​x−zm)‖ℒ1
≤ ‖gk⋆(𝟏𝕂​f)−hm−1g‖ℒ1​(𝕂k)
+‖gk⋆(𝟏𝕂​f)−hm−1g‖ℒ1​(𝕂k∁)
+‖cm​kmn​g​(km​x−zm)‖ℒ1
≤ ϵ8+cm+‖gk⋆(𝟏𝕂​f)‖ℒ1​(𝕂k∁)+‖hm−1g‖ℒ1​(𝕂k∁)​,

where ‖cm​kmn​g​(km​x−zm)‖ℒ1≤cm since kmn​g​(km​x−zm) is a pdf. The aim is now to prove that

‖gk⋆(𝟏𝕂​f)‖ℒ1​(𝕂k∁)​<ϵ24​ and ∥​hm−1g∥ℒ1​(𝕂k∁)<ϵ24​.

Using polar coordinates and (17), we have

∫𝟏{x:‖x−y‖2>k−γ}​gk​(x−y)​d​λ​(x)
≤∫𝟏{x:‖x−y‖2>k−γ}​β​k−θ(k−1+‖x−y‖2)n+θ​d​λ​(x)
=β​An​k−θ​∫𝟏(k−γ,∞)​rn−1(k−1+r)n+θ​d​λ​(r)
≤β​An​k−θ​∫𝟏(k−γ,∞)​r−θ−1​d​λ​(r)
=β​An​kθ​(γ−1)/θ​,

where An is the surface area of a unit sphere embedded in ℝn. We then have

‖gk⋆(𝟏𝕂​f)‖ℒ1​(𝕂k∁)
=∫∫𝟏{y∈𝕂}​𝟏{x∈𝕂k∁}​f​(y)​gk​(x−y)​d​λ​(x)​d​λ​(y)
≤‖𝟏𝕂​f‖ℒ∞​∫∫𝟏{y∈𝕂}​𝟏{x:‖x−y‖2>k−γ}​β​k−θ(k−1+‖x−y‖2)n+θ​d​λ​(x)​d​λ​(y)
≤‖𝟏𝕂​f‖ℒ∞​λ​(𝕂)​β​An​kθ​(γ−1)/θ​,

which implies that we can choose a k3∈ℕ, such that for all k≥k3,

‖gk⋆(𝟏𝕂​f)‖ℒ1​(𝕂k∁)<ϵ24​. (22)

Lastly, we write

‖hm−1g‖ℒ1​(𝕂k∁)
=∫𝟏𝕂k∁​∑i=1m−1ci​kn​g​(k​x−zi)​d​λ
=∑i=1m−1∫𝟏𝕂k∁​[k−n​∫𝟏𝔸i​f​(zk)​d​λ​(z)]​kn​g​(k​x−zi)​d​λ​(x)
≤‖𝟏𝕂​f‖ℒ∞​∑i=1m−1k−n​λ​(𝔸i)​∫𝟏𝕂k∁​gk​(x−zik)​d​λ
≤‖𝟏𝕂​f‖ℒ∞​∑i=1m−1k−n​λ​(𝔸i)​β​An​kθ​(γ−1)θ
≤‖𝟏𝕂​f‖ℒ∞​λ​(𝕂)​β​An​kθ​(γ−1)/θ​,

which implies that we can choose the same k3 as above to obtain the bound

‖hm−1g‖ℒ1​(𝕂k∁)<ϵ24​, (23)

for any k≥k3.

Thus, we obtain the bound ‖𝟏𝕂​f−hmg‖ℒ1<ϵ/2, for all k≥max⁡{k1,k2,k3}, by combining (18), (19), (20), (21), (22), and (23), via the triangle inequality. The result is proved by combing the bound above, with (16), for an appropriately large r​(ϵ)∈ℕ.

3 Comments and discussion

3.1 Relationship to Theorem 1

In the proof of Theorem 1, the famous Hilbert space approximation result of [16] and [2] was used to bound the ℒ2 norm between any approximand f∈ℒ2 and a convex combination of bounded functions in ℒ2. This approximation theorem is exactly the p=2 case of the more general theorem of [9], as presented in Lemma 11. Thus, one can view Theorem 5(c) as the p∈(1,∞) generalization of Theorem 1.

3.2 The class 𝒲 is a proper subset of the class 𝒞0

Here, we comment on the nature of class 𝒲, which was investigated by [1] and [26]. We recall that [1] conjectured that Theorem 4 generalizes from g=ϕ to g∈𝒱. In Theorem 5(a)–(e), we assume that g∈𝒞0. We can demonstrate that g∈𝒞0 is a strictly weaker condition than g∈𝒱 or g∈𝒲.

For example, consider the function in g:ℝ→ℝ such that g​(x)=0 if x<0 and

g​(x) =∑i=1∞22​ii​[(x−i+1)2​i​𝟏{i−1≤x<i−1/2}+(x−i)2​i​𝟏{i−1/2≤x<i}]​ if ​x≥0​,

and note that

∫𝟏(−1/2,1/2)​(2​x)2​ii​d​λ=12​i2+i<1i2​.

Since ∑i=1∞(1/i2)=π2/6, g∈ℒ1. Furthermore, g is continuous since all stationary points of g are continuous. In ℝ, g∈𝒞0 if

limx→±∞g​(x)=0​.

For x≤0, we observe that g=0 and thus the left limit is satisfied. On the right, for any 1/ϵ>0, we have x​(ϵ)≥⌈ϵ⌉−1/2, so that g​(x)<1/ϵ, for all x>x​(ϵ), where ⌈⋅⌉ is the ceiling operator. Therefore, g∈𝒞0.

Within each interval i−1≤x<i, we observe that g is locally maximized at x=i−1/2. The local maximum corresponding to each of these points is 1/i. Thus g∉𝒲, since

∑i=1∞1i​<∑y∈ℤsupx∈[0,1] |​g​(x+y)|,

where ∑i=1∞(1/i)=∞. Furthermore, g∉𝒱 since 𝒱⊂𝒲.

3.3 Convergence in measure

Along with the conclusions of Theorem 5(d) and (e), Lemma 12 also implies convergence in measure. That is, if ν is a σ​-finite Borel measure on ℝn, then for any ν​-measurable f, there exists a sequence {hmg}, such that for any ϵ>0,

limm→∞υ​({x∈ℝn:|f​(x)−hmg​(x)|≥ϵ})=0​.

Appendix A Technical results

Throughout the main text, we utilize a number of established technical results. For the convenience of the reader, we append these results within this Appendix. Sources from which we draw the unproved results are provided at the end of the section.

Lemma 5.

Let {gk} be a sequence of pdfs in ℒ1 and for every δ>0

limk→∞∫𝟏{x:‖x‖2>δ}​gk​d​λ=0​.

Then, for all f∈ℒp and 1≤p<∞,

limk→∞‖gk⋆f−f‖ℒp=0​.

Furthermore, for all f∈𝒞b and any compact 𝕂⊂ℝn,

limk→∞‖gk⋆f−f‖ℒ∞​(𝕂)=0​.

The sequences {gk} from Lemma 5 are often called approximate identities or approximations of the identity. A simple construction of approximate identities is by taking dilations gk​(x)=kn​g​(k​x), which yields the following corollary.

Corollary 1.

Let g be a pdf. Then the sequence of dilations {gk:gk​(x)=kn​g​(k​x)}, satisfies the hypothesis of Lemma 5 and hence permits its conclusion.

Lemma 6.

The class 𝒞0 is a subset of 𝒞b. Furthermore, if f∈𝒞0, then f is uniformly continuous.

Lemma 7 (Urysohn’s Lemma).

If 𝕂⊂ℝn is compact, then there exists some g∈𝒞c, such that 0≤g≤1 and 𝟏𝕂​g=1.

Lemma 8.

If 𝕏⊂ℝn is bounded, then for any r>0, 𝕏 can be covered by ⋃i=1m𝔹​(xi,r) for some finite m∈ℕ, where xi∈ℝn and i∈[m].

Lemma 9.

If 0<p<q<r≤∞, then ℒp∩ℒr⊂ℒq.

Let Γ:ℝ→ℝ be the usual gamma function, defined as Γ​(z)=∫𝟏(0,∞)​xz−1​exp⁡(−x)​d​λ.

Lemma 10.

If f∈ℒp and g∈ℒ1, for 1≤p≤∞, then f⋆g exists and we have ‖f⋆g‖ℒp≤‖g‖ℒ1​‖f‖ℒp.

Lemma 11.

Let 𝒢⊂ℒp, for some 1≤p<∞, and let f∈Conv¯​(𝒢). For any K>0, such that ‖f−α‖ℒp<K, for all α∈𝒢, there exists a hm∈Convm​(𝒢), such that

‖f−hm‖ℒp≤Cp​Km1−1/α​,

where α=min⁡{p,2}, and

Cp={1if ​1≤p≤2​,2​[π​Γ​(p+12)]1/pif ​p>2​.
Lemma 12.

In any measure ν, uniform convergence implies almost uniform convergence, and almost uniform convergence implies almost everywhere convergence and convergence in measure, with respect to ν.

Appendix B Sources of results

Lemma 5 is reported as Theorem 9.3.3 in [21] (see also Theorem 2 of [4, Ch. 20]). The proof of Corollary 1 can be taken from that of Theorem 4 of [4, Ch. 20]. Lemma 6 appears in [5], as Proposition 1.4.5. Lemma 7 is taken from Corollary 1.2.9 of [5]. Lemma 8 appears as Theorem 1.2.2 in [5]. Lemma 9 can be found in [12, Prop. 6.10]. Lemma 10 can be found in [21, Thm. 9.3.1]. Lemma 11 appears as Corollary 2.6 in [9]. Lemma 12 can be obtained from the definition of almost uniform convergence, Lemma 7.10, and Theorem 7.11 of [3].

Acknowledgment

HDN is personally funded by Australian Research Council (ARC) grant DE170101134. HDN and GJM are supported by ARC grant DP180101192. FC is supported by Agence Nationale de la Recherche (ANR) grant SMILES ANR-18-CE40-0014 and by Région Normandie grant RIN AStERiCs.

References

  • [1] A G Bacharoglou. Approximation of probability distributions by convex mixtures of Gaussian measures. Proceedings of the American Mathematical Society, 138:2619–2628, 2010. ↩ 1 3.2
  • [2] A R Barron. Universal approximation bound for superpositions of a sigmoidal function. IEEE Transactions on Information Theory, 39:930–945, 1993. ↩ 2.4 3.1
  • [3] R G Bartle. The Elements of Integration and Lesbegue Measure. Wiley, New York, 1995. ↩ B
  • [4] W Cheney and W Light. A Course in Approximation Theory. Brooks/Cole, Pacific Grove, 2000. ↩ 1 2.1 B
  • [5] J B Conway. A Course in Abstract Analysis. American Mathematical Society, Providence, 2012. ↩ B
  • [6] A DasGupta. Asymptotic Theory Of Statistics And Probability. Springer, New York, 2008. ↩ 1
  • [7] L De Branges. The Stone-Weierstrass theorem. Proceedings of the American Mathematical Society, 10:822–824, 1959. ↩ 1
  • [8] L Devroye and G Lugosi. Combinatorial Methods in Density Estimation. Springer, New York, 2000. ↩ 1
  • [9] M J Donahue, L Gurvits, C Darken, and E Sontag. Rates of convex approximation in non-Hilbert spaces. Constructive Approximation, 13:187–220, 1997. ↩ 1 2.4 3.1 B
  • [10] B S Everitt and D J Hand. Finite Mixture Distributions. Chapman and Hall, London, 1981. ↩ 1
  • [11] H G Feichtinger. A characterization of Wiener’s algebra on locally compact groups. Archiv der Mathematik, 29:136–140, 1977. ↩ 1
  • [12] G B Folland. Real Analysis: Modern Techniques and Their Applications. Wiley, New York, 1999. ↩ B
  • [13] F Forbes and D Wraith. A new family of multivariate heavy-tailed distributions with variable marginal amounts of tailweights. Statistics and Computing, 24:971–984, 2013. ↩ 1
  • [14] S Fruwirth-Schnatter. Finite Mixture and Markov Switching Models. Springer, New York, 2006. ↩ 1
  • [15] S Fruwirth-Schnatter, G Celeux, and C P Robert, editors. Handbook of Mixture Analysis. CRC Press, Boca Raton, 2019. ↩ 1
  • [16] L K Jones. A simple lemma on greedy approximation in Hilbert space and convergene rates for projection pursuit regression and neural network training. Annals of Statistics, 20:608–613, 1992. ↩ 2.4 3.1
  • [17] S Kullback and R A Leibler. On information and sufficiency. Annals of Mathematical Statistics, 22:79–86, 1951. ↩ 1
  • [18] S X Lee and G J McLachlan. Finite mixtures of canonical fundamental skew t-distributions: the unification of the trstricted and unrestricted skew t-mixture models. Statistics and Computing, 26:573–589, 2016. ↩ 1
  • [19] J Q Li and A R Barron. Mixture density estimation. In S A Solla, T K Leen, and K R Mueller, editors, Advances in Neural Information Processing Systems, volume 12, Cambridge, 1999. MIT Press. ↩ 1
  • [20] B G Lindsay. Mixture models: theory, geometry and applications. In NSF-CBMS Regional Conference Series in Probability and Statistics, 1995. ↩ 1
  • [21] B Makarov and A Podkorytov. Real Analysis: Measures, Integrals and Applications. Springer, London, 2013. ↩ 2.1 B
  • [22] G J McLachlan and K E Basford. Mixture Models: Inference and Applications to Clustering. Marcel Dekker, New York, 1988. ↩ 1
  • [23] G J McLachlan and D Peel. Finite Mixture Models. Wiley, New York, 2000. ↩ 1
  • [24] Kerrie L Mengersen, Christian Robert, and Mike Titterington. Mixtures: Estimation and Applications. Wiley, New York, 2011. ↩ 1
  • [25] V Nestoridis and C Papadimitropoulos. Abstract theory of universal series and an application to Dirichlet series. Comptes Rendus Academy of Science Paris Series I, 341:530–543, 2005. ↩ 1
  • [26] V Nestoridis, S Schmutzhard, and V Stefanopoulos. Universal series induced by approximate identities and some relevant applications. Journal of Approximation Theory, 163:1783–1797, 2011. ↩ 1 2.1 3.2
  • [27] V Nestoridis and V Stefanopoulos. Universial series and approximate identities. Technical report, Department of Mathematics and Statistics, University of Cyprus, 2007. ↩ 1 2.1
  • [28] H D Nguyen and G J McLachlan. On approximations via convolution-defined mixture models. Communications in Statistics - Theory and Methods, 2019. In press. ↩ 1
  • [29] D Peel and G J McLachlan. Robust mixture modelling using the t distribution. Statistics and Computing, 10:339–348, 2000. ↩ 1
  • [30] A Rakhlin, D Panchenko, and S Mukherjee. Risk bounds for mixture density estimation. ESAIM: Probability and Statistics, 9:220–229, 2005. ↩ 1
  • [31] W Rudin. Principles of Mathematical Analysis. McGraw-Hill, Singapore, 1976. ↩ 2.1
  • [32] I W Sandberg. Gaussian radial basis functions and inner product sspace. Circuits, Systems and Signal Processing, 20:635–642, 2001. ↩ 1
  • [33] P Schlattmann. Medical Applications of Finite Mixture Models. Springer, Berlin, 2009. ↩ 1
  • [34] M H Stone. The generalized Weierstrass approximation theorem. Mathematical Magazine, 21:237–254, 1948. ↩ 1
  • [35] D M Titterington, A F M Smith, and U E Makov. Statistical Analysis of Finite Mixture Distributions. Wiley, New York, 1985. ↩ 1
  • [36] S van de Geer. Asymptotic theory for maximum likelihood in nonparametric mixture models. Computational Statistics and Data Analysis, 41:453–464, 2003. ↩ 2.4
  • [37] J L Walker and M Ben-Akiva. Advances in discrete choice: mixture models. In Andre De Palma, Robin Lindsey, Emile Quinet, and Roger Vickerman, editors, A Handbook of Transport Economics, pages 160–187. Edward Edgar, Cheltenham, 2011. ↩ 1
  • [38] Y Xu, W A Light, and E W Cheney. Constructive methods of approximation by ridge functions and radial functions. Numerical Algorithms, 4:205–223, 1993. ↩ 1
  • [39] G Yona. Introduction to Computational Proteomics. CRC Press, Boca Raton, 2011. ↩ 1
  • [40] A J Zeevi and R Meir. Density estimation through convex combinations of densities: approximation and estimation bounds. Neural Computation, 10:99–109, 1997. ↩ 1

Cite this paper

Please cite the published version. Venue: Cogent Mathematics & Statistics, Open-access journal article (Vol. 7, 2020). DOI: 10.1080/25742558.2020.1750861. Official record: Cogent Math. Stat..

BibTeX
@article{nguyen2020approximation,
  title     = {Approximation by finite mixtures of continuous density functions that vanish at infinity},
  author    = {{TrungTin Nguyen} and {Hien D. Nguyen} and {Faicel Chamroukhi} and {Geoffrey J. McLachlan}},
  journal   = {Cogent Mathematics \& Statistics},
  volume    = {7}, number = {1}, pages = {1750861},
  year      = {2020}, publisher = {Taylor \& Francis},
  doi       = {10.1080/25742558.2020.1750861},
}