Publications · Journal

Approximations of conditional probability density functions in Lebesgue spaces via mixture of experts models

Hien D. Nguyen†, TrungTin Nguyen, Faicel Chamroukhi, Geoffrey J. McLachlan

† Corresponding author.

J. Stat. Distrib. Appl. · Journal Journal of Statistical Distributions and Applications. Open-access journal article (Vol. 8, Art. 13, 2021).

Abstract

Mixture of experts (MoE) models are widely applied for conditional probability density estimation problems. We demonstrate the richness of the class of MoE models by proving denseness results in Lebesgue spaces, when inputs and outputs variables are both compactly supported. We further prove an almost uniform convergence result when the input is univariate. Auxiliary lemmas are proved regarding the richness of the soft-max gating function class, and their relationships to the class of Gaussian gating functions.

1Department of Mathematics and Statistics, La Trobe University, Bundoora Victoria, Australia.
2Normandie Univ, UNICAEN, CNRS, LMNO, 14000 Caen, France.
3School of Mathematics and Physics, University of Queensland, St. Lucia Brisbane, Australia.
∗Corresponding author—email: h.nguyen5@latrobe.edu.au.

Key words: mixture of experts; conditional probability density functions; approximation theory; mixture models; Lebesgue spaces

1 Introduction

Mixture of experts (MoE) models are a widely applicable class of conditional probability density approximations that have been considered as solution methods across the spectrum of statistical and machine learning problems; see, for example, the reviews of Yuksel et al., (2012), Masoudnia & Ebrahimpour, (2014), and Nguyen & Chamroukhi, (2018).

Let ℤ=𝕏×𝕐, where 𝕏⊆ℝd and 𝕐⊆ℝq, for d,q∈ℕ. Suppose that the input and output random variables, 𝑿∈𝕏 and 𝒀∈𝕐, are related via the conditional probability density function (PDF) f​(𝒚|𝒙) in the functional class:

ℱ={f:ℤ→[0,∞)|∫𝕐f(𝒚|𝒙)dλ(𝒚)=1,∀𝒙∈𝕏},

where λ denotes the Lebesgue measure. The MoE approach seeks to approximate the unknown target conditional PDF f by a function of the MoE form:

m​(𝒚|𝒙)=∑k=1KGatek​(𝒙)​Expertk​(𝒚)​,

where 𝐆𝐚𝐭𝐞=(Gatek)k∈[K]∈𝒢K ([K]={1,…,K}), Expert1,…,ExpertK∈ℰ, and K∈ℕ. Here, we say that m is a K​-component MoE model with gates arising from the class 𝒢K and experts arising from ℰ, where ℰ is a class of PDFs with support 𝕐.

The most popular choices for 𝒢K are the parametric soft-max and Gaussian gating classes:

𝒢SK={𝐆𝐚𝐭𝐞=(Gatek​(⋅;𝜸))k∈[K]|∀k∈[K],Gatek​(⋅;𝜸)=exp(ak+𝒃k⊤⋅)∑l=1Kexp(al+𝒃l⊤⋅),𝜸∈𝔾SK}

and

𝒢GK={𝐆𝐚𝐭𝐞=(Gatek​(⋅;𝜸))k∈[K]|∀k∈[K],Gatek​(⋅;𝜸)=πk​ϕ​(⋅;𝝂k,𝚺k)∑l=1Kπl​ϕ​(⋅;𝝂l,𝚺l),𝜸∈𝔾GK}​,

respectively, where

𝔾SK={𝜸=(a1,…,aK,𝒃1,…,𝒃K)∈ℝK×(ℝd)K}

and

𝔾GK={𝜸=(𝝅,𝝂1,…,𝝂K,𝚺1,…,𝚺K)∈ΠK−1×(ℝd)K×𝕊dK}​.

Here,

ϕ(⋅;𝝂,𝚺)=|2π𝚺|−1/2exp[−12(⋅−𝝂)⊤𝚺−1(⋅−𝝂)]

is the multivariate normal density function with mean vector 𝝂 and covariance matrix 𝚺, 𝝅⊤=(π1,…,πK) is a vector of weights in the simplex:

ΠK−1={𝝅=(πk)k∈[K]|∀k∈[K],πk>0,∑k=1Kπk=1}​,

and 𝕊d is the class of d×d symmetric positive definite matrices. The soft-max and Gaussian gating classes were first introduced by Jacobs et al., (1991) and Jordan & Xu, (1995), respectively. Typically, one chooses experts that arise from some location-scale class:

ℰψ={gψ​(⋅;𝝁,σ):𝕐→[0,∞)|gψ​(⋅;𝝁,σ)=1σq​ψ​(⋅−𝝁σ),𝝁∈ℝq,σ∈(0,∞)}​,

where ψ is a PDF, with respect to ℝq in the sense that ψ:ℝq→[0,∞) and ∫ℝqψ​(𝒚)​d​λ​(𝒚)=1.

We shall say that f∈ℒp​(ℤ) for any p∈[1,∞) if

‖f‖p,ℤ=(∫ℤ|𝟏ℤ​f|p​d​λ​(𝒛))1/p<∞​,

where 𝟏ℤ is the indicator function that takes value 1 when 𝒛∈ℤ, and 0 otherwise. Further, we say that f∈ℒ∞​(ℤ) if

‖f‖∞,ℤ=inf{a≥0|λ​({𝒛∈ℤ||f​(𝒛)|>a})=0}<∞​.

We shall refer to ∥⋅∥p,ℤ as the ℒp norm on ℤ, for p∈[0,∞], and where the context is obvious, we shall drop the reference to ℤ.

Suppose that the target conditional PDF f is in the class ℱp=ℱ∩ℒp. We address the problem of approximating f, with respect to the ℒp norm, using MoE models in the soft-max and Gaussian gated classes,

ℳSψ = {mKψ:ℤ→[0,∞)|mKψ(𝒚|𝒙)=∑k=1KGatek(𝒙)gψ(𝒚;𝝁k,σk),
gψ∈ℰψ∩ℒ∞,𝐆𝐚𝐭𝐞∈𝒢SK,𝝁k∈𝕐,σk∈(0,∞),k∈[K],K∈ℕ}

and

ℳGψ = {mKψ:ℤ→[0,∞)|mKψ(𝒚|𝒙)=∑k=1KGatek(𝒙)gψ(𝒚;𝝁k,σk),
gψ∈ℰψ∩ℒ∞,𝐆𝐚𝐭𝐞∈𝒢GK,𝝁k∈𝕐,σk∈(0,∞),k∈[K],K∈ℕ},

by showing that both ℳSψ and ℳGψ are dense in the class ℱp, when 𝕏=[0,1]d and 𝕐 is a compact subset of ℝq. Our denseness results are enabled by the indicator function approximation result of Jiang & Tanner, 1999b , and the finite mixture model denseness theorems of Nguyen et al., 2020a and Nguyen et al., 2020c .

Our theorems contribute to an enduring continuity of sustained interest in the approximation capabilities of MoE models. Related to our results are contributions regarding the approximation capabilities of the conditional expectation function of the classes ℳSψ and ℳGψ (Wang & Mendel, 1992, Zeevi et al., 1998, Jiang & Tanner, 1999b, Krzyzak & Schafer, 2005, Mendes & Jiang, 2012, Nguyen et al., 2016, 2019) and the approximation capabilities of subclasses of ℳSψ and ℳGψ, with respect to the Kullback–Leibler divergence (Jiang & Tanner, 1999a, Norets, 2010, Norets & Pelenis, 2014). Our results can be seen as complements to the Kullback–Leibler approximation theorems of Norets, (2010) and Norets & Pelenis, (2014), by the relationship between the Kullback–Leibler divergence and the ℒ2 norm (Zeevi & Meir, 1997). That is, when f>1/κ, for all (𝒙,𝒚)∈ℤ and some constant κ>0, we have that the integrated conditional Kullback–Leibler divergence considered by Norets & Pelenis, (2014):

∫𝕏D(f(⋅|𝒙)∥mKψ(⋅|𝒙))dλ(𝒙)=∫𝕏∫𝕐f(𝒚|𝒙)logf​(𝒚|𝒙)mKψ​(𝒚|𝒙)dλ(𝒚)dλ(𝒙)

satisfies

∫𝕏D(f(⋅|𝒙)∥mKψ(⋅|𝒙))dλ(𝒙)≤κ2∥f−mKψ∥2,ℤ2,

and thus a good approximation in the integrated Kullback–Leibler divergence is guaranteed if one can find a good approximation in the ℒ2 norm, which is guaranteed by our main result.

The remainder of the manuscript proceeds as follows. The main result is presented in Section 2. Technical lemmas are provided in Section 3. The proofs of our results are then presented in Section 4. Proofs of required lemmas that do not appear elsewhere are provided in Section 5. A summary of our work and some conclusions are drawn in Section 6.

2 Main results

Denote the class of bounded functions on ℤ by

ℬ​(ℤ)={f∈ℒ∞​(ℤ)|∃a∈[0,∞)​, such that ​|f​(𝒛)|≤a​, ​∀𝒛∈ℤ}​,

and write its norm as ‖f‖ℬ​(ℤ)=sup𝒛∈ℤ|f​(𝒛)|. Further, let 𝒞 denote the class of continuous functions. Note that if ℤ is compact and f∈𝒞, then f∈ℬ.

Theorem 1.

Assume that 𝕏=[0,1]d for d∈ℕ. There exists a sequence {mKψ}K∈ℕ⊂ℳSψ, such that if 𝕐⊂ℝq is compact, f∈ℱ∩𝒞, and ψ∈𝒞​(ℝq) is a PDF on support ℝq, then limK→∞‖f−mKψ‖p=0, for p∈[1,∞) .

Since convergence in Lebesgue spaces does not imply point-wise modes of convergence, the following result is also useful and interesting in some restricted scenarios. Here, we note that the mode of convergence is almost uniform, which implies almost everywhere convergence and convergence in measure (cf. Bartle, 1995, Lem 7.10 and Thm. 7.11). The almost uniform convergence of {mKψ}K∈ℕ to f in the following result is to be understood in the sense of Bartle, (1995, Def. 7.9). That is, for every δ>0, there exists a set 𝔼δ⊂ℤ with λ​(ℤ)<δ, such that {mKψ}K∈ℕ converges to f, uniformly on ℤ\𝔼δ.

Theorem 2.

Assume that 𝕏=[0,1]. There exists a sequence {mKψ}K∈ℕ⊂ℳSψ, such that if 𝕐⊂ℝq is compact, f∈ℱ∩𝒞, and ψ∈𝒞​(ℝq) is a PDF on support ℝq, then limK→∞mKψ=f, almost uniformly.

The following result establishes the connection between the gating classes 𝒢SK and 𝒢GK.

Lemma 1.

For each K∈ℕ, 𝒢SK⊂𝒢GK. Further, if we define the class of Gaussian gating vectors with equal covariance matrices:

𝒢EK={𝐆𝐚𝐭𝐞=(Gatek​(⋅;𝜸))k∈[K]|∀k∈[K],Gatek​(⋅;𝜸)=πk​ϕ​(⋅;𝝂k,𝚺)∑l=1Kπl​ϕ​(⋅;𝝂l,𝚺),𝜸∈𝔾EK}​,

where

𝔾EK={𝜸=(𝝅,𝝂1,…,𝝂K,𝚺)∈ΠK−1×(ℝd)K×𝕊d}​,

then 𝒢EK⊂𝒢SK.

We can directly apply Lemma 1 to establish the following corollary to Theorems 1 and 2, regarding the approximation capability of the class ℳGψ.

Corollary 1.

Theorems 1 and 2 hold when ℳSψ is replaced by ℳGψ in their statements.

3 Technical lemmas

Let 𝕂n={(k1,…,kd)∈[n]d} and κ:𝕂n→[nd] be a bijection for each n∈ℕ. For each (k1,…,kd)∈𝕂n and k∈[nd], we define 𝕏kn=𝕏κ​(k1,…,kd)n=∏i=1d𝕀kin, where 𝕀kin=[(ki−1)/n,ki/n) for ki∈[n−1], and 𝕀nn=[(n−1)/n,1].

We call {𝕏kn}k∈[nd] a fine partition of 𝕏, in the sense that 𝕏=[0,1]d=⋃k=1nd𝕏kn, for each n, and that λ​(𝕏kn)=n−d gets smaller, as n increases. The following result from Jiang & Tanner, 1999b establishes the approximation capability of soft-max gates.

Lemma 2 (Jiang and Tanner, 1999, p. 1189).

For each n∈ℕ, p∈[1,∞) and ϵ>0, there exists a gating functions

𝐆𝐚𝐭𝐞=(Gatek​(⋅;𝜸))k∈[nd]∈𝒢Snd

for some 𝛄∈𝔾Snd, such that

supk∈[nd]‖𝟏{𝒙∈𝕏kn}−Gatek​(⋅;𝜸)‖p,𝕏≤ϵ​.

When, d=1, we have also the following almost uniform convergence alternative to Lemma 2.

Lemma 3.

Let 𝕏=[0,1]. Then, for each n∈ℕ, there exists a sequence of gating functions:

{𝐆𝐚𝐭𝐞l=(Gatek​(⋅;𝜸l))k∈[nd]}l∈ℕ⊂𝒢Sn​,

defined by {𝛄l}l∈ℕ⊂𝔾Sn, such that

Gatek​(⋅;𝜸l)→𝟏{𝒙∈𝕏kn}​,

almost uniformly, simultaneously for all k∈[nd].

For PDF ψ on support ℝq, define the class of finite mixture models by

ℋψ = {hKψ:ℝq→[0,∞)|hKψ(𝒚)=∑k=1Kckgψ(𝒚;𝝁k,σk),
gψ∈ℰψ∩ℒ∞,(ck)k∈[K]∈ΠK−1,𝝁k∈𝕐,σk∈(0,∞),k∈[K],K∈ℕ}.

We require the following result, from Nguyen et al., 2020a , regarding the approximation capabilities of ℋψ.

Lemma 4 (Nguyen et al., 2020a, Thm. 2(b)).

If f∈𝒞​(𝕐) is a PDF on 𝕐, ψ∈𝒞​(ℝq) is a PDF on ℝq, and 𝕐⊂ℝq is compact, then there exists a sequence {hKψ}K∈ℕ⊂ℋψ, such that limK→∞‖f−hKψ‖ℬ​(𝕐)=0.

4 Proofs of main results

4.1 Proof of Theorem 1

To prove the result, it suffices to show that for each ϵ>0, there exists a mKψ∈ℳSψ, such that

‖f−mKψ‖p<ϵ​.

The main steps of the proof are as follows. We firstly approximate f​(𝒚|𝒙) by

υn​(𝒚|𝒙)=∑k=1nd𝟏{𝒙∈𝕏kn}​f​(𝒚|𝒙kn)​, (1)

where 𝒙kn∈𝕏kn, for each k∈[nd], such that

‖f−υn‖p<ϵ3​, (2)

for all n≥N1​(ϵ), for some sufficiently large N1​(ϵ)∈ℕ. Then we approximate υn​(𝒚|𝒙) by

ηn​(𝒚|𝒙)=∑k=1ndGatek​(𝒙;𝜸n)​f​(𝒚|𝒙kn)​, (3)

where 𝜸n∈𝔾Snd and 𝐆𝐚𝐭𝐞=(Gatek​(⋅;𝜸n))k∈[nd]∈𝒢Snd, so that

‖υn−ηn‖p ≤supk∈[nd]∥Gatek(⋅;𝜸)−𝟏{𝒙∈𝕏kn}∥p,𝕏∑k=1nd∥f(⋅|𝒙kn)∥p,𝕐<ϵ3, (4)

using Lemma 2.

Finally, we approximate ηn​(𝒚|𝒙) by mKnψ​(𝒚|𝒙), where

mKnψ​(𝒚|𝒙)=∑k=1ndGatek​(𝒙;𝜸)​hnkk​(𝒚|𝒙kn) (5)

and

hnkk​(𝒚|𝒙kn)=∑i=1nkcik​gψ​(𝒚;𝝁ik,σik)∈ℋψ (6)

for nk∈ℕ (k∈[nd]), such that Kn=∑k=1ndnk. Here, we establish that there exists N2​(ϵ,n,𝜸n)∈ℕ, so that when nk≥N2​(ϵ,n,𝜸n),

∥ηn−mKnψ∥p≤supk∈[nd]∥Gatek(⋅;𝜸)∥p,𝕏∑k=1nd∥f(⋅|𝒙kn)−hnkk(⋅|𝒙kn)∥p,𝕐<ϵ3. (7)

Results (2)–(7) then imply that for each ϵ>0, there exists N1​(ϵ), 𝜸n, and N2​(ϵ,n,𝜸n), such that for all Kn=∑k=1ndnk, where nk≥N2​(ϵ,n,𝜸n) (for each k∈[nd]) and n≥N1​(ϵ). The following inequality results from an application of the triangle inequality:

‖f−mKnψ‖p ≤‖f−υn‖p+‖υn−ηn‖p+‖ηn−mKnψ‖p<3×ϵ3=ϵ​.

We now focus our attention to proving each of the results: (2)–(7). To prove (2), we note that since f is uniformly continuous (because ℤ=𝕏×𝕐 is compact, and f∈𝒞), there exists a function (1) such that for all ε>0,

sup(𝒙,𝒚)∈ℤ|f(𝒚|𝒙)−υ(𝒚|𝒙)|<ε. (8)

We can construct such an approximation by considering the fact that as n increases, the diameter δn=supk∈nddiam​(𝕏kn) of the fine partition goes to zero. By the uniform continuity of f, for every ε>0, there exists a δ​(ϵ)>0, such that if ‖(𝒙1,𝒚1)−(𝒙2,𝒚2)‖<δ​(ϵ), then |f(𝒚1|𝒙1)−f(𝒚2|𝒙2)|<ε, for all pairs (𝒙1,𝒚1),(𝒙2,𝒚2)∈ℤ. Here, ∥⋅∥ denotes the Euclidean norm. Furthermore, for any (𝒙,𝒚)∈ℤ, we have

|f(𝒚|𝒙)−υn(𝒚|𝒙)| =|∑k=1nd𝟏{𝒙∈𝕏kn}[f(𝒚|𝒙)−f(𝒚|𝒙kn)]|
≤∑k=1nd𝟏{𝒙∈𝕏kn}|f(𝒚|𝒙)−f(𝒚|𝒙kn)|, (9)

by the triangle inequality.

Since 𝒙kn∈𝕏kn, for each k and n, we have the fact that ‖(𝒙,𝒚)−(𝒙kn,𝒚)‖<δn for (𝒙,𝒚)∈𝕏kn×𝕐. By uniform continuity, for each ε, we can find a sufficiently small δ​(ϵ), such that |f(𝒚|𝒙)−f(𝒚|𝒙kn)|<ε, if ‖(𝒙,𝒚)−(𝒙kn,𝒚)‖<δ​(ϵ), for all k. The desired result (8) can be obtained by noting that the right hand side of (9) consists of only one non-zero summand for any (𝒙,𝒚)∈ℤ, and by choosing n∈ℕ sufficiently large, so that δn<δ​(ϵ).

By (8), we have the fact that υn→f, point-wise. We can bound υn as follows:

υn​(𝒚|𝒙)≤∑i=1np𝟏{𝒙∈𝕏kn}​sup𝜻∈𝕐,𝝃∈𝕏f​(𝜻|𝝃)=sup𝜻∈𝕐,𝝃∈𝕏f​(𝜻|𝝃)​, (10)

where the right-hand side is a constant and is therefore in ℒp, since ℤ is compact. An application of the Lebesgue dominated convergence theorem in ℒp then yields (2).

Next we write

‖υn−ηn‖p =∥∑k=1nd𝟏{𝒙∈𝕏kn}f(𝒚|𝒙kn)−∑k=1ndGatek(𝒙;𝜸n)f(𝒚|𝒙kn)∥p
≤∑k=1nd∥[𝟏{𝒙∈𝕏kn}−Gatek(𝒙;𝜸n)]f(𝒚|𝒙kn)∥p.

Since the norm arguments are separable in 𝒙 and 𝒚, we apply Fubini’s theorem to get

‖υn−ηn‖p =∑k=1nd∥[𝟏{𝒙∈𝕏kn}−Gatek(𝒙;𝜸n)]∥p,𝕏∥f(𝒚|𝒙kn)∥p,𝕐
≤supk∈[nd]∥[𝟏{𝒙∈𝕏kn}−Gatek(𝒙;𝜸n)]∥p,𝕏∑k=1nd∥f(𝒚|𝒙kn)∥p,𝕐

Because f∈ℬ and nd is finite, for any fixed n∈ℕ, we have C1(n)=∑k=1nd∥f(𝒚|𝒙kn)∥p,𝕐<∞. For each ϵ>0, we need to choose a 𝜸n∈𝔾Snd, such that

supk∈[nd]‖[𝟏{𝒙∈𝕏kn}−Gatek​(𝒙;𝜸n)]‖p,𝕏<ϵ3​C1​(n)​,

which can be achieved via a direct application of Lemma 2. We have thus shown (4).

Lastly, we are required to approximate f​(𝒚|𝒙kn) for each k∈[nd], by a function of form (6). Since 𝕐 is compact and f and ψ are continuous, we can apply of Lemma 4, directly. Note that over a set of finite measure, convergence in ∥⋅∥ℬ implies convergence in ℒp norm, for all p∈[1,∞] (cf. Oden & Demkowicz, 2010, Prop. 3.9.3).

We can then write (5) as

mKnψ​(𝒚|𝒙) =∑k=1ndexp⁡(an,k+𝒃n,k⊤​𝒙)∑l=1ndexp⁡(an,l+𝒃n,l⊤​𝒙)​hnkk​(𝒚|𝒙kn)
=∑k=1nd∑i=1nkexp⁡(an,k+𝒃n,k⊤​𝒙)∑l=1ndexp⁡(an,l+𝒃n,l⊤​𝒙)​cik∑l=1nkclk​gψ​(𝒚;𝝁ik,σik)
=∑k=1nd∑i=1nkexp⁡(log⁡cik+an,k+𝒃n,k⊤​𝒙)∑l=1nd∑j=1nkexp⁡(log⁡cjk+an,l+𝒃n,l⊤​𝒙)​gψ​(𝒚;𝝁ik,σik)​, (11)

where 𝜸n=(an,1,…,an,nd,𝒃n,1,…,𝒃n,nd). From (11), we observe that mKnψ∈ℳSψ, with Kn=∑k=1ndnk.

To obtain (7), we write

‖ηn−mKnψ‖p =∥∑k=1ndGatek(𝒙;𝜸n)f(𝒚|𝒙kn)−∑k=1ndGatek(𝒙;𝜸)hnkk(𝒚|𝒙kn)∥p
≤∑k=1nd∥Gatek(𝒙;𝜸n)[f(𝒚|𝒙kn)−hnkk(𝒚|𝒙kn)]∥p.

By separability and Fubini’s theorem, we then have

‖ηn−mKnψ‖ ≤∑k=1nd∥Gatek(𝒙;𝜸n)∥p,𝕏∥f(𝒚|𝒙kn)−hnkk(𝒚|𝒙kn)∥p,𝕐
≤supk∈[nd]∥Gatek(𝒙;𝜸n)∥p,𝕏∑k=1nd∥f(𝒚|𝒙kn)−hnkk(𝒚|𝒙kn)∥p,𝕐.

Let C2​(n,𝜸n)=supk∈[nd]‖Gatek​(𝒙;𝜸n)‖p,𝕏. Then, we apply Lemma 4 nd times to establish the existence of a constant N2​(ϵ,n,𝜸n)∈ℕ, such that for all k∈[nd] and nk≥N2​(ϵ,n,𝜸n),

∥f(𝒚|𝒙kn)−hnkk(𝒚|𝒙kn)∥p,𝕐≤ϵ3​C2​(n,𝜸n)​nd.

Thus, we have

‖ηn−mKnψ‖≤C2​(n,𝜸n)×nd×ϵ3​C2​(n,𝜸n)​nd=ϵ3​,

which completes our proof.

4.2 Proof of Theorem 2

The proof is procedurally similar to that of Theorem 1 and thus we only seek to highlight the important differences. Firstly, for any ϵ>0, we approximate f​(𝒚|𝒙) by υn​(𝒙|𝒚) of form (1), with d=1. Result (2) implies uniform convergence, in the sense that there exists an N1​(ϵ)∈ℕ, such that for all n≥N1​(ϵ),

‖f−υn‖ℬ<ϵ3​. (12)

We now seek to approximate υn by ηn of form (3), with 𝜸n=𝜸l for some l∈ℕ. Upon application of Lemma 3, it follows that for each k∈[nd] and ε>0, there exists a measurable set 𝔹k​(ε)⊆𝕏, such that

λ​(𝔹k​(ε))<εnd​λ​(𝕐)

and

‖Gatek​(⋅;𝜸l)−𝟏{𝒙∈𝕏kn}‖ℬ​(𝔹k∁​(ε))<ϵ3

for all l≥Mk​(ϵ,n), for some Mk​(ϵ,n)∈ℕ. Here, (⋅)∁ is the set complement operator.

Since f∈ℬ, we have the bound C(n)=∑k=1nd∥f(𝒚|𝒙kn)∥ℬ​(𝕐)<∞. Write 𝔹​(ε)=⋃k=1nd𝔹k​(ε). Then, 𝔹∁​(ε)=⋂k=1nd𝔹k∁​(ε),

λ​(𝔹​(ε))≤∑k=1ndλ​(𝔹k∁​(ε))<ελ​(𝕐)​,

and

‖Gatek​(⋅;𝜸l)−𝟏{𝒙∈𝕏kn}‖ℬ​(𝔹∁​(ε)) ≤mink∈[nd]⁡‖Gatek​(⋅;𝜸l)−𝟏{𝒙∈𝕏kn}‖ℬ​(𝔹k∁​(ε))<ϵ3​C​(n)​,

for all l≥M​(ϵ,n)=maxk∈[nd]⁡Mk​(ϵ,n). Here we use the fact that the supremum over some intersect of sets is less than or equal to the minimum of the supremum over each individual set.

Upon defining ℂ​(ε)=𝔹​(ε)×𝕐⊂ℤ, we observe that

λ​(ℂ​(ε))=λ​(𝔹​(ε))​λ​(𝕐)≤ελ​(𝕐)×λ​(𝕐)=ε​,

and ℂ​(ε)⊂𝔹​(ε)×𝕐. Note also that

(𝔹​(ε)×𝕐)∁=ℤ\(𝔹​(ε)×𝕐)=𝔹∁​(ε)×𝕐

and

ℂ∁​(ε)=(𝔹∁​(ε)×𝕐)∪(𝔹​(ε)×𝕐∁)∪(𝔹∁​(ε)×𝕐∁)​.

It follows that

‖υn−ηn‖ℬ​(ℂ∁​(ε))≤max⁡{‖υn−ηn‖ℬ​(𝔹∁​(ε)×𝕐),‖υn−ηn‖ℬ​(𝔹​(ε)×𝕐∁),‖υn−ηn‖ℬ​(𝔹∁​(ε)×𝕐∁)}​.

Since 𝔹​(ε)×𝕐∁ and 𝔹∁​(ε)×𝕐∁ are empty, via separability, we have

‖υn−ηn‖ℬ​(ℂ∁​(ε)) =‖υn−ηn‖ℬ​(𝔹∁​(ε)×𝕐)
=sup𝒛∈𝔹∁​(ε)×𝕐|∑k=1nd[𝟏{𝒙∈𝕏kn}−Gatek(𝒙;𝜸l)]f(𝒚|𝒙kn)|
≤sup𝒛∈𝔹∁​(ε)×𝕐∑k=1nd|𝟏{𝒙∈𝕏kn}−Gatek​(𝒙;𝜸l)|​f​(𝒚|𝒙kn)
≤∑k=1ndsup𝒛∈𝔹∁​(ε)×𝕐|𝟏{𝒙∈𝕏kn}−Gatek​(𝒙;𝜸l)|​f​(𝒚|𝒙kn)
=∑k=1nd∥𝟏{𝒙∈𝕏kn}−Gatek(𝒙;𝜸l)∥ℬ​(𝔹∁​(ε))∥f(𝒚|𝒙kn)∥ℬ​(𝕐)
≤supk∈[n]∥𝟏{𝒙∈𝕏kn}−Gatek(𝒙;𝜸l)∥ℬ​(𝔹∁​(ε))∑k=1nd∥f(𝒚|𝒙kn)∥ℬ​(𝕐).

Recall that the ∑k=1nd∥f(𝒚|𝒙kn)∥ℬ​(𝕐)=C(n)<∞ and that we can choose l≥M​(ϵ,n) so that

supk∈[n]‖𝟏{𝒙∈𝕏kn}−Gatek​(𝒙;𝜸l)‖ℬ​(𝔹∁​(ε))<ϵ3​C​(n)​,

and thus

‖υn−ηn‖ℬ​(ℂ∁​(ε))<ϵ3​C​(n)×C​(n)=ϵ3​, (13)

as required.

Finally, by noting that for each k∈[nd], both (6) and f(⋅|𝒙kn) are continuous over 𝕐, we apply Lemma 4 to obtain an N2​(ϵ,n,l)∈ℕ, such that for any ϵ>0 and nk≥N2​(ϵ,n,l), we have

∥f(⋅|𝒙kn)−hnkk(⋅|𝒙kn)∥ℬ​(𝕐)<ϵ3​M1​n.

Here M1=supk∈[nd]‖Gatek​(⋅;𝜸l)‖ℬ​(𝕏)<∞, since Gatek​(𝒙;𝜸l) is continuous in 𝒙, and 𝕏 is compact. Therefore, for all Kn=∑k=1ndnk, nk≥N2​(ϵ,n,l),

‖ηn−mKnψ‖ℬ ≤supk∈[nd]∥Gatek(𝒙;𝜸l)∥ℬ​(𝕏)∑k=1nd∥f(⋅|𝒙kn)−hnkk(⋅|𝒙kn)∥ℬ​(𝕐)
=M1×nd×ϵ3​M1​nd=ϵ3. (14)

In summary, via (12), (13), and (14), for each ϵ>0, for any ε>0, there exists a ℂ​(ε)⊂ℤ and constants N1​(ϵ),M​(ϵ,n),N2​(ϵ,n,l)∈ℕ, such that for all Kn=∑k=1ndnk, with nk≥N2​(ϵ,n,l), l≥M​(ϵ,n), and n≥N1​(ϵ), it follows that λ​(ℂ​(ε))<ε, and

‖f−mKnψ‖ℬ​(ℂ∁​(ε)) ≤‖f−υn‖ℬ​(ℂ∁​(ε))+‖υn−ηn‖ℬ​(ℂ∁​(ε))+‖ηn−mKnψ‖ℬ​(ℂ∁​(ε))
≤‖f−υn‖ℬ+‖υn−ηn‖ℬ​(ℂ∁​(ε))+‖ηn−mKnψ‖ℬ
<3×ϵ3=ϵ​.

This completes the proof.

5 Proofs of lemmas

5.1 Proof of Lemma 1

We firstly prove that any gating vector from 𝒢SK can be equivalently represented as an element of 𝒢GK. For any 𝒙∈ℝd, d∈ℕ, k∈[K], ak∈ℝ, 𝒃k∈ℝd, and K∈ℕ, choose 𝝂k=𝒃k, τk=ak+𝒃k⊤​𝒃k/2 and

πk=exp⁡(τk)/∑l=1Kexp⁡(τl)​.

This implies that ∑l=1Kπl=1, πl>0, for all l∈[K], and

exp⁡(ak+𝒃k⊤​𝒙)∑l=1Kexp⁡(ak+𝒃k⊤​𝒙) =exp⁡(τk−𝒗k⊤​𝒗k/2+𝒗k⊤​𝒙)∑l=1Kexp⁡(τl−𝒗l⊤​𝒗l/2+𝒗l⊤​𝒙)
=exp⁡(τk)​exp⁡(−(𝒙−𝝂k)⊤​(𝒙−𝝂k)/2)∑l=1Kexp⁡(τl)​exp⁡(−(𝒙−𝝂l)⊤​(𝒙−𝝂l)/2)
=πk​(2​π)−d/2​exp⁡(−(𝒙−𝝂k)⊤​(𝒙−𝝂k)/2)∑l=1Kπl​(2​π)−d/2​exp⁡(−(𝒙−𝝂l)⊤​(𝒙−𝝂l)/2)
=πk​ϕ​(𝒙;𝝂k,𝐈)∑l=1Kπl​ϕ​(𝒙;𝝂l,𝐈)​,

where 𝐈 is the identity matrix of appropriate size. This proves that 𝒢SK⊂𝒢GK.

Next, to show that 𝒢EK⊂𝒢SK, we write

πk​ϕ​(𝒙;𝝂k,𝚺)∑l=1Kπl​ϕ​(𝒙;𝝂l,𝚺)
= πk​|2​π​𝚺|−1/2​exp⁡(−(𝒙−𝝂k)⊤​𝚺−1​(𝒙−𝝂k)/2)∑l=1Kπl​|2​π​𝚺|−1/2​exp⁡(−(𝒙−𝝂l)⊤​𝚺−1​(𝒙−𝝂l)/2)
= 1∑l=1Kexp⁡(−log⁡(πl−2/πk−2)/2−(𝒙−𝝂l)⊤​𝚺−1​(𝒙−𝝂l)/2−(𝒙−𝝂k)⊤​𝚺−1​(𝒙−𝝂k)/2)​,

and note that

(𝒙−𝝂l)⊤​𝚺−1​(𝒙−𝝂l)−(𝒙−𝝂k)⊤​𝚺−1​(𝒙−𝝂k)
= −2​(𝝂l−𝝂k)⊤​𝚺−1​𝒙+(𝝂l+𝝂k)⊤​𝚺−1​(𝝂l−𝝂k)​.

Thus, we have

πk​ϕ​(𝒙;𝝂k,𝚺)∑l=1Kπl​ϕ​(𝒙;𝝂l,𝚺)
= 1∑l=1Kexp⁡(−log⁡(πl−2/πk−2)/2−(𝝂l+𝝂k)⊤​𝚺−1​(𝝂l−𝝂k)/2−(𝝂l−𝝂k)⊤​𝚺−1​𝒙)​.

Next, notice that we can write

exp⁡(ak+𝒃k⊤​𝒙)∑l=1Kexp⁡(al+𝒃l⊤​𝒙)=1∑l=1Kexp⁡(αl+𝜷l⊤​𝒙)​,

where αl=al−ak and 𝜷l=𝜷l−𝜷k. We now choose ak and 𝒃k, such that for every l∈[K],

αl =al−ak=−12​log⁡(πl−2πk−2)−12​(𝝂l⊤​𝚺−1​𝝂l−𝝂k⊤​𝚺−1​𝝂k)​,

and

𝜷l=𝜷l−𝜷k=𝝂l⊤​𝚺−1−𝝂k⊤​𝚺−1​.

To complete the proof, we choose

ak=log⁡(πk)−12​𝝂k⊤​𝚺−1​𝝂k

and bk=𝝂k⊤​𝚺−1, for each k∈[K].

5.2 Proof of Lemma 3

For l∈[0,∞), write

Gatek​(x,l)=exp⁡([x−ck]​l​k)∑i=1nexp⁡([x−ci]​l​i)​,

where x∈𝕏=[0,1], and ck=(k−1)/(2​k). We identify that 𝐆𝐚𝐭𝐞=(Gatek​(x,l))k∈[n] belongs to the class 𝒢Sn. The proof of the Section 4 Proposition from Jiang & Tanner, 1999b reveals that for all k∈[n],

Gatek​(x,l)→𝟏{x∈𝕀kn}

almost everywhere in λ, as l→∞. The result then follows via an application of Egorov’s theorem (cf. Folland, 1999, Thm. 2.33).

6 Summary and conclusions

Using recent results mixture model approximation results Nguyen et al., 2020a and Nguyen et al., 2020c , and the indicator approximation theorem of Jiang & Tanner, 1999b (cf. Section 3), we have proved two approximation theorems (Theorems 1 and 2) regarding the class of soft-max gated MoE models with experts arising from arbitrary location-scale families of conditional density functions. Via an equivalence result (Lemma 1), the results of Theorems 1 and 2 also extend to the setting of Gaussian gated MoE models (Corollary 1), which can be seen as a generalization of the soft-max gated MoE models.

Although we explicitly make the assumption that 𝕏=[0,1]d, for the sake of mathematical argument (so that we can make direct use of Lemma 2), a simple shift-and-scale argument can be used to generalize our result to cases where 𝕏 is any generic compact domain. The compactness assumption regarding the input domain is common in the MoE and mixture of regression models literature, as per the works of Jiang & Tanner, 1999a , Jiang & Tanner, 1999b , Norets, (2010), Montuelle & Le Pennec, (2014), Pelenis, (2014), Devijver, 2015a , and Devijver, 2015b .

The assumption permits the application of the result to the settings where the inputs 𝑿 is assumed to be non-random design vectors that take value on some compact set 𝕏. This is often the case when there is only a finite number of possible design vector elements for which 𝑿 can take. Otherwise, the assumption also permits the scenario where 𝑿 is some random element with compactly supported distribution, such as uniformly distributed, or beta distributed inputs. Unfortunately, the case of random 𝑿 over an unbounded domain (e.g., if 𝑿 has multivariate Gaussian distribution) is not covered under our framework. An extension to such cases would require a more general version of Lemma 2, which we believe is a nontrivial direction for future work.

Like the input, we also assume that the output domain is restricted to a compact set 𝕐. However, the output domain of the approximating class of MoE models is unrestricted to 𝕐 and thus the functions (i.e., we allow ψ to be a PDF over ℝq). The restrictions placed on 𝕐 is also common in the mixture approximation literature, as per the works of Zeevi & Meir, (1997), Li & Barron, (1999), and Rakhlin et al., (2005), and is also often made in the context of nonparametric regression (see, e.g., Gyorfi et al., 2002 and Cucker & Zhou, 2007). Here, our use of the compactness of 𝕐 is to bound the integral of vn, in (10). A more nuanced approach, such as via the use of a generalized Lebesgue spaces (see e.g., Castillo & Rafeiro, 2010 and Cruze-Uribe & Fiorenza, 2013), may lead to result for unbounded 𝕐. This is another exciting future direction of our research program.

A trivial modification to the proof of Lemma 4 allows us to replace the assumption that f is a PDF with a sub-PDF assumption (i.e., ∫𝕐f​d​λ≤1), instead. This in turn permits us to replace the assumption that f(⋅|𝒙) is a conditional PDF in Theorems 1 and 2 with sub-PDF assumptions as well (i.e., for each 𝒙∈𝕏, ∫𝕐f​(𝒚|𝒙)​d​λ​(𝒚)≤1). Thus, in this modified form, we have a useful interpretation for situations when the input 𝒀 is unbounded. That is, when 𝒀 is unbounded, we can say that the conditional PDF f can be arbitrarily well approximated in ℒp norm by a sequence {mKψ}K∈ℕ of either soft-max or Gaussian gated MoEs over any compact subdomain 𝕐 of the unbounded domain of 𝒀. Thus, although we cannot provide guarantees of the entire domain of 𝒀, we are able to guarantee arbitrary approximate fidelity over any arbitrarily large compact subdomain. This is a useful result in practice since one is often not interested in the entire domain of 𝒀, but only on some subdomain where the probability of 𝒀 is concentrated. This version of the result resembles traditional denseness results in approximation theory, such as those of Cheney & Light, (2000, Ch. 20).

Finally, our results can be directly applied to provide approximation guarantees for a large number of currently used models in applied statistics and machine learning research. Particularly, our approximation guarantees are applicable to the recent MoE models of Ingrassia et al., (2012), Chamroukhi et al., (2013), Ingrassia et al., (2014), Chamroukhi, (2016), Nguyen & McLachlan, (2016), Deleforge et al., 2015b , Deleforge et al., 2015a , Kalliovirta et al., (2016), and Perthame et al., (2018), among many others. Here, we may guarantee that the underlying data generating processes, if satisfying our assumptions, can be adequately well approximated by sufficiently complex forms of the models considered in each of the aforementioned work.

The rate and manner of which good approximation can be achieved as a function of the number of experts K and the sample size is a currently active research area, with pioneering work conducted in Cohen & Le Pennec, (2012) and Montuelle & Le Pennec, (2014). More recent results in this direction appear in Nguyen et al., 2020b , Nguyen et al., 2021b , and Nguyen et al., 2021a .

List of abbreviations

MoE

Mixture of experts

PDF

Probability density function

Acknowledgements

Hien Duy Nguyen and Geoffrey John McLachlan are funded by Australian Research Council grants: DP180101192 and IC170100035. TrungTin Nguyen is supported by a “Contrat doctoral” from the French Ministry of Higher Educationand Research. Faicel Chamroukhi is granted by the French National Research Agency (ANR) grant SMILES ANR-18-CE40-0014. The authors also thank the Editor and Reviewer, whose careful and considerate comments lead to improvements in of text.

Declarations

Funding: HDN and GJM are funded by Australian Research Council grants: DP180101192 and IC170100035. FC is funded by ANR grant: SMILES ANR-18-CE40-0014.
Conflicts of interest: None.
Availability of data and material: Not applicable.
Authors’ contributions: All authors contributed equally to the exposition and to the mathematical derivations.
Code availability: Not applicable.

References

  • Bartle, (1995) Bartle, R. (1995). The Elements of Integration and Lebesgue Measure. New York: Wiley. ↩ 2
  • Castillo & Rafeiro, (2010) Castillo, R. E. & Rafeiro, H. (2010). An Introductory Course in Lebesgue Spaces. Switzerland: Springer. ↩ 6
  • Chamroukhi, (2016) Chamroukhi, F. (2016). Robust mixture of experts modeling using the t distribution. Neural Networks, 79, 20–36. ↩ 6
  • Chamroukhi et al., (2013) Chamroukhi, F., Mohammed, S., Trabelsi, D., Oukhellou, L., & Amirat, Y. (2013). Joint segmentation of multivariate time series with hidden process regression for human activity recognition. Neurocomputing, 120, 633–644. ↩ 6
  • Cheney & Light, (2000) Cheney, W. & Light, W. (2000). A Course in Approximation Theory. Pacific Grove: Brooks/Cole. ↩ 6
  • Cohen & Le Pennec, (2012) Cohen, S. & Le Pennec, E. (2012). Conditional density estimation by penalized likelihood model selection and application. ArXiv, (arXiv:1103.2021). ↩ 6
  • Cruze-Uribe & Fiorenza, (2013) Cruze-Uribe, D. V. & Fiorenza, A. (2013). Variable Lebesgue Spaces: Foundations and Harmonic Analysis. Basel: Birkhauser. ↩ 6
  • Cucker & Zhou, (2007) Cucker, F. & Zhou, D.-X. (2007). Learning Theory: An Approximation Theory Viewpoint. Cambridge: Cambridge University Press. ↩ 6
  • (9) Deleforge, A., Forbes, F., & Horaud, R. (2015a). Acoustic space learning for sound-source separation and localization on binaural manifolds. International Journal of Neural Systems, 25, 1440003. ↩ 6
  • (10) Deleforge, A., Forbes, F., & Horaud, R. (2015b). High-dimensional regression with Gaussian mixtures and partially-latent response variables. Statistics and Computing, 25, 893–911. ↩ 6
  • (11) Devijver, E. (2015a). An ℓ1-oracle inequality for the Lasso in multivariate finite mixture of multivariate Gaussian regression models. ESAIM: Probability and Statistics, 19, 649–670. ↩ 6
  • (12) Devijver, E. (2015b). Finite mixture regression: a sparse variable selection by model selection for clustering. Electronic Journal of Statistics, 9, 2642–2674. ↩ 6
  • Folland, (1999) Folland, G. B. (1999). Real Analysis: Modern Techniques and Their Applications. New York: Wiley. ↩ 5.2
  • Gyorfi et al., (2002) Gyorfi, L., Kohler, M., Krzyzak, A., & Walk, H. (2002). A Distribution-free Theory Of Nonparametric Regression. New York: Springer. ↩ 6
  • Ingrassia et al., (2014) Ingrassia, S., Minotti, S. C., & Punzo, A. (2014). Model-based clustering via linear cluster-weighted models. Computational Statistics and Data Analysis, 71, 159–182. ↩ 6
  • Ingrassia et al., (2012) Ingrassia, S., Minotti, S. C., & Vittadini, G. (2012). Local statistical modeling via a cluster-weighted approach with elliptical distributions. Journal of Classification, 29, 363–401. ↩ 6
  • Jacobs et al., (1991) Jacobs, R. A., Jordan, M. I., Nowlan, S. J., & Hinton, G. E. (1991). Adaptive mixtures of local experts. Neural Computation, 3, 79–87. ↩ 1
  • (18) Jiang, W. & Tanner, M. A. (1999a). Hierachical mixtures-of-experts for exponential family regression models: approximation and maximum likelihood estimation. Annals of Statistics, 27, 987–1011. ↩ 1 6
  • (19) Jiang, W. & Tanner, M. A. (1999b). On the approximation rate of hierachical mixtures-of-experts for generalized linear models. Neural Computation, 11, 1183–1198. ↩ 1 3 5.2 6
  • Jordan & Xu, (1995) Jordan, M. I. & Xu, L. (1995). Convergence results for the EM approach to mixtures of experts architectures. Neural Networks, 8, 1409–1431. ↩ 1
  • Kalliovirta et al., (2016) Kalliovirta, L., Meitz, M., & Saikkonen, P. (2016). Gaussian mixture vector autoregression. Journal of Econometrics, 192, 485–498. ↩ 6
  • Krzyzak & Schafer, (2005) Krzyzak, A. & Schafer, D. (2005). Nonparametric regression estimation by normalized radial basis function networks. IEEE Transactions on Information Theory, 51, 1003–1010. ↩ 1
  • Li & Barron, (1999) Li, J. Q. & Barron, A. R. (1999). Mixture density estimation. In S. A. Solla, T. K. Leen, & K. R. Mueller (Eds.), Advances in Neural Information Processing Systems, volume 12 Cambridge: MIT Press. ↩ 6
  • Masoudnia & Ebrahimpour, (2014) Masoudnia, S. & Ebrahimpour, R. (2014). Mixture of experts: a literature survey. Artificial Intelligence Review, 42, 275–293. ↩ 1
  • Mendes & Jiang, (2012) Mendes, E. F. & Jiang, W. (2012). On convergence rates of mixture of polynomial experts. Neural Computation, 24. 3025-3051. ↩ 1
  • Montuelle & Le Pennec, (2014) Montuelle, L. & Le Pennec, E. (2014). Mixture of Gaussian regressions model with logistic weights, a penalized maximum likelihood approach. Electronic Journal of Statistics, 8, 1661–1695. ↩ 6
  • Nguyen & Chamroukhi, (2018) Nguyen, H. D. & Chamroukhi, F. (2018). Practical and theoretical aspects of mixture-of-experts modeling: an overview. WIREs Data Mining and Knowledge Discovery, (pp. e1246). ↩ 1
  • Nguyen et al., (2019) Nguyen, H. D., Chamroukhi, F., & Forbes, F. (2019). Approximation results regarding the multiple-output Gaussian gated mixture of linear experts model. Neurocomputing, 366, 208–214. ↩ 1
  • Nguyen et al., (2016) Nguyen, H. D., Lloyd-Jones, L. R., & McLachlan, G. J. (2016). A universal approximation theorem for mixture-of-experts models. Neural Computation, 28, 2585–2593. ↩ 1
  • Nguyen & McLachlan, (2016) Nguyen, H. D. & McLachlan, G. J. (2016). Laplace mixture of linear experts. Computational Statistics and Data Analysis, 93, 177–191. ↩ 6
  • (31) Nguyen, T. T., Chamroukhi, F., Nguyen, H. D., & Forbes, F. (2021a). Non-asymptotic model selection in block-diagonal mixture of polynomial experts models. arXiv preprint arXiv:2104.08959. ↩ 6
  • (32) Nguyen, T. T., Chamroukhi, F., Nguyen, H. D., & McLachlan, G. J. (2020a). Approximation of probability density functions via location-scale finite mixtures in Lebesgue spaces. arXiv:2008.09787. ↩ 1 3 6
  • (33) Nguyen, T. T., Nguyen, H. D., Chamroukhi, F., & Forbes, F. (2021b). A non-asymptotic penalization criterion for model selection in mixture of experts models. ArXiv, (arXiv:2104.02640). ↩ 6
  • (34) Nguyen, T. T., Nguyen, H. D., Chamroukhi, F., & McLachlan, G. J. (2020b). An l1-oracle inequality for the Lasso in mixture-of-experts regression models. ArXiv, (arXiv:2009.10622). ↩ 6
  • (35) Nguyen, T. T., Nguyen, H. D., Chamroukhi, F., & McLachlan, G. J. (2020c). Approximation by finite mixtures of continuous density functions that vanish at infinity. Cogent Mathematics and Statistics, 7, 1750861. ↩ 1 6
  • Norets, (2010) Norets, A. (2010). Approximation of conditional densities by smooth mixtures of regressions. Annals of Statistics, 38, 1733–1766. ↩ 1 6
  • Norets & Pelenis, (2014) Norets, A. & Pelenis, J. (2014). Posterior consistency in conditional density estimation by covariate dependent mixtures. Econometric Theory, 30, 606–646. ↩ 1
  • Oden & Demkowicz, (2010) Oden, J. T. & Demkowicz, L. F. (2010). Applied Functional Analysis. Boca Raton: CRC Press. ↩ 4.1
  • Pelenis, (2014) Pelenis, J. (2014). Bayesian regression with heteroscedastic error density and parametric mean function. Journal of Econometrics, 178, 624–638. ↩ 6
  • Perthame et al., (2018) Perthame, E., Forbes, F., & Deleforge, A. (2018). Inverse regression approach to robust nonlinear high-to-low dimensional mapping. Journal of Multivariate Analysis, 163, 1–14. ↩ 6
  • Rakhlin et al., (2005) Rakhlin, A., Panchenko, D., & Mukherjee, S. (2005). Risk bounds for mixture density estimation. ESAIM: Probability and Statistics, 9, 220–229. ↩ 6
  • Wang & Mendel, (1992) Wang, L.-X. & Mendel, J. M. (1992). Fuzzy basis functions, universal approximation, and orthogonal least-squares learning. IEEE Transactions on Neural Networks, 3, 807–814. ↩ 1
  • Yuksel et al., (2012) Yuksel, S. E., Wilson, J. N., & Gader, P. D. (2012). Twenty years of mixture of experts. IEEE Transactions on Neural Networks and Learning Systems, 23, 1177–1193. ↩ 1
  • Zeevi & Meir, (1997) Zeevi, A. J. & Meir, R. (1997). Density estimation through convex combinations of densities: approximation and estimation bounds. Neural Computation, 10, 99–109. ↩ 1 6
  • Zeevi et al., (1998) Zeevi, A. J., Meir, R., & Maiorov, V. (1998). Error bounds for functional approximation and estimation using mixtures of experts. IEEE Transactions on Information Theory, 44, 1010–1025. ↩ 1

Cite this paper

Please cite the published version. Venue: Journal of Statistical Distributions and Applications, Open-access journal article (Vol. 8, Art. 13, 2021). DOI: 10.1186/s40488-021-00125-0. Official record: SpringerOpen.

BibTeX
@article{nguyen2021approximations,
  title     = {Approximations of conditional probability density functions in Lebesgue spaces via mixture of experts models},
  author    = {Nguyen, Hien D. and Nguyen, TrungTin and Chamroukhi, Faicel and McLachlan, Geoffrey J.},
  journal   = {Journal of Statistical Distributions and Applications},
  volume    = {8}, number = {1}, pages = {13},
  year      = {2021}, publisher = {Springer},
  doi       = {10.1186/s40488-021-00125-0},
}