Software Demos · Interactive demo

The Asymptotic Law of the Minimum Empirical Risk

Reshape a risk valley from a sharp point to a flat plateau and watch the limit law of the minimum empirical risk morph from a clean Gaussian into the skewed infimum of a Gaussian process, the over-specification regime where naive Gaussian inference and the bootstrap fail. Use the buttons beside each control (or the Run the experiments launchers) to auto-play; everything runs client-side.

The Asymptotic Law of the Minimum Empirical Risk

Fit a model by minimizing empirical risk and the smallest attainable risk, the minimum empirical risk (MER), is a random number. Its rescaled fluctuation √n (MER − min f) has a limit law you can use for inference: bootstrap confidence sets, goodness-of-fit tests, model selection. Mode 1 (MER limit law) shows how that limit morphs from a clean Gaussian (unique minimizer) into a skewed Gaussian-process infimum (over-specified flat plateau). Mode 2 (MoE model selection) is the paper's dependent-data (non-i.i.d.) application: choose the number of experts K in a softmax-gated Gaussian Mixture-of-Experts, where the MER's √n rate dictates a dependence-corrected penalty (SWIC) that recovers the true K₀, while the naive i.i.d. BIC penalty over-selects.

Empirical risk minimizationAsymptotic distributionOver-specificationGaussian-process infimumMixture-of-experts model selectionDependent data
Minimum empirical risk: two simulations
teal = population risk f · blue = one empirical risk f̂ₙ = f + n−1/2G · teal band = argmin set S · amber = this draw's minimizer
histogram of √n (MER − min f) · teal curve = classical N(0,σ²)

The unknown is a hypothesis x ∈ [0,1]. The population risk f(x) (teal) is a valley whose bottom you reshape: at w = 0 it is a sharp parabola with a single minimizer x₀ = ½; as w grows the bottom flattens into a plateau, so the population argmin set S = [x₀ − w/2, x₀ + w/2] becomes an interval (teal band). Each draw is a fresh empirical risk f̂ₙ(x) = f(x) + n−1/2 G(x) (blue), where G is a smooth mean-zero Gaussian field (a random-Fourier surrogate for the empirical process, correlation length , marginal variance σ²). We locate its minimizer (amber), read off the minimum empirical risk, and drop the rescaled blow-up √n (MER − min f) = infx [√n f(x) + G(x)] into the histogram. Singleton (w = 0): the minimizer localizes at x₀ and the statistic converges to G(x₀) ~ N(0, σ²), symmetric and centered, and the teal Gaussian overlay fits, so the CLT and the bootstrap are valid. Plateau (w > 0): the statistic becomes infx∈S G(x), the infimum of a Gaussian process over S, which is negatively shifted and left-skewed, so the histogram detaches from the Gaussian: this is the over-specification regime where naive Gaussian inference and the standard bootstrap fail. Raising n tightens the singleton toward the Gaussian but leaves the plateau's non-Gaussian law unchanged. (Honest toy: this is the i.i.d., continuous, √n baseline with a random-Fourier Gaussian-field stand-in for the empirical process and a grid argmin. The paper proves the general limit for non-i.i.d. data, discontinuous losses and losses indexed by non-Euclidean spaces, then builds consistent bootstrap confidence sets, penalized model-selection tests, and neural-network applications: none of that heavier machinery is in this 2D cartoon.)

Run the experiments

Every animation runs live in your browser. Click a button to run that experiment on the demo (it scrolls up and starts, switching to the right simulation mode); drag any control to take over. Nothing is downloaded, it is generated on the fly.

Mode 1 · MER limit law

Widen the plateau

Sweep the flat-bottom half-width from a sharp unique minimizer (w = 0) to a broad argmin set S, and watch the histogram of the rescaled minimum empirical risk detach from the Gaussian, shift left and skew.

Mode 1 · MER limit law

Roughen the noise field

Change the correlation length of the Gaussian field driving the empirical risk: a rougher field gives the infimum over the plateau more chances to dip, deepening the left shift and skew.

Mode 1 · MER limit law

Grow the sample size

Increase n and see the singleton case tighten toward the exact N(0, sigma^2), while the over-specified plateau keeps its non-Gaussian, skewed limit no matter how large n gets.

Mode 1 · MER limit law

Singleton vs plateau

Toggle between the unique-minimizer (identifiable) case, whose histogram matches the symmetric Gaussian, and the flat-plateau (over-specified) case, whose limit is the skewed infimum of a Gaussian process where the bootstrap fails.

Mode 1 · MER limit law

Accumulate fresh draws

Drop many fresh empirical-risk draws into the histogram one at a time to trace out the limiting law of the rescaled minimum empirical risk for the current valley shape.

Mode 2 · MoE selection

Turn up the dependence

Increase the order m of the m-dependent input sequence, shrinking the effective sample size while the true experts K0 stay fixed. Watch each criterion curve and its selected K respond as the inputs cluster into correlated runs.

Mode 2 · MoE selection

Change the truth K0

Switch the true number of softmax-gated experts. At K0 = 2 the SWIC histogram sits exactly on K0 and AIC collapses to 1; push K0 = 3 to expose the penalty tradeoff, where the heavy SWIC penalty can under-select at modest n while the weak BIC over-selects.

Mode 2 · MoE selection

Grow the sample

Increase the number of dependent observations and watch every criterion sharpen, with SWIC locking onto K0 fastest.

Mode 2 · MoE selection

Accumulate the selection histogram

Run many independent trials, each generating fresh dependent SGaME data and fitting K = 1..Kmax by EM, to build the paper's histogram of the selected model for SWIC, BIC and AIC.

The idea in three steps

Empirical risk minimization returns both an estimate and a number, the minimum empirical risk. That number fluctuates, and its limit law is the engine behind bootstrap confidence sets, goodness-of-fit tests, and the model-selection penalties that pick the right model complexity, provided you know which regime you are in.

1 · Rescale

The √n blow-up

The gap between the minimum empirical risk and the population minimum is tiny, of order 1/n. Multiply by √n and it converges to a genuine random variable: infx[√n f(x) + G(x)], driven by a Gaussian process G from the CLT for the risk.

2 · Identify

One point, one Gaussian

If the population risk has a unique minimizer, the √n term pins the infimum to that point and the limit is just G there: a symmetric N(0,σ²). Here the CLT and the bootstrap deliver valid confidence sets and tests.

3 · Select

A rate that sets the penalty

Over-specified models (a whole flat argmin set S) give a shifted, skewed Gaussian-process-infimum limit. Under dependent data this √n rate is what a consistent model-selection penalty must match: the SWIC penalty df log n/√n recovers the true number of experts K₀ where the naive i.i.d. BIC over-selects.

For the general limit theorems under non-i.i.d. data, discontinuous and non-Euclidean-indexed losses, the consistent bootstrap confidence sets, penalized model-selection tests and the mixture-of-experts experiments, see On the Asymptotic Distribution of the Minimum Empirical Risk (Westerhout, TrungTin Nguyen, Guo & H. D. Nguyen, ICML 2024, PMLR Vol. 235). The mixture-of-experts simulation reproduces NIIDMoE/BIC_NIID_SGaME.R using the meteorits softmax-gated normal MoE.