Software Demos · Interactive demo

Robust Bayesian Parameter Balancing

Balancing the thermodynamic parameters of a reaction system so the constraints hold, robustly, under heavy-tailed and skewed measurement error. Use the buttons beside each control (or the Run the experiments launchers) to auto-play; everything runs client-side.

Robust Bayesian Parameter Balancing

Reconcile noisy, inconsistent kinetic measurements into thermodynamically consistent parameters with calibrated uncertainty, and watch robust likelihoods shrug off outliers.

Bayesian inferenceUncertainty quantificationSystems biologyRobustness
Phosphoglucose isomerase (G6P ⇌ F6P) · Haldane balancing
drag a measurement ↕ · click a track to add / remove data · log scale
Coverage of 95% credible intervals
covered
calibrated uncertainty covers the truth ~95% of the time; an overconfident fit does not.
Haldane consistency (|Δ ln Keq|)
raw  →  balanced 0.00
the balanced estimate satisfies the Haldane relation exactly; the raw measurements do not.
Haldane relation · PGI (G6P ⇌ F6P)
ln Keq = ln kcat⁺ − ln kcat⁻
       + ln KM(F6P) − ln KM(G6P)

Kinetic parameters curated from databases (kcat and KM from BRENDA / SABIO-RK, Keq from eQuilibrator) are noisy and mutually inconsistent. Bayesian parameter balancing treats them as noisy observations of the free log-parameters and enforces the thermodynamic relations, so the posterior is consistent by construction with calibrated credible intervals. In PGI mode the Haldane relation ties Keq to kcat and KM. In Pathway mode the shared metabolite F6P couples the two reactions two ways: its formation energy links the equilibria Keq₁, Keq₂, and its concentration links the reaction quotients Q₁, Q₂, so dragging any measurement propagates through the shared quantity to the other reaction (switch a track off and watch it get inferred). Under the Gaussian likelihood, add an outlier or turn on Skewed errors and the estimates bias upward (for outliers, coverage collapses); the Student-t resists heavy tails, and the skew-normal additionally recovers the true values under one-sided, right-skewed noise. (Values are illustrative; the constraints are the real thermodynamic ones.)

Run the experiments

Every animation runs live in your browser. Click a button to play that experiment on the demo (it scrolls up and starts); drag the slider to take over. Nothing is downloaded.

Robustness to outliers

Add measurement outliers one by one; the robust posterior barely moves while the Gaussian one is dragged off.

Gaussian / Student-t / Skew-normal

Tour the three noise models and compare their bias and interval coverage under skewed, heavy-tailed errors.

Measurement noise

Sweep the measurement-noise level and watch the balanced estimate and its credible interval respond.

Sweep source heterogeneity (figures)

Scroll to Figures 1 & 2 and sweep the source-heterogeneity τ: the random-effects bars widen selectively while the posterior means stay put.

Tour the parameter blocks

Cycle the three observable blocks, standard chemical potentials, kcat and KM, of the all-models-vs-database figure.

Consensus vs sources

Raise the between-source heterogeneity in the random-effects panel and watch the fixed-effects interval fail while the RE interval keeps covering the consensus.

Watch calibration break

Sweep the error misspecification: the Gaussian coverage falls below nominal (overconfident) while the robust and random-effects intervals stay honest.

Figure 5: random-effects lift

On the paper's real leave-one-source-out coverage curves (raw BRENDA glycolysis), turn on source-level random effects and watch all three families rise from under-covering toward the nominal diagonal.

Scaling: exact vs structured VI

Tour the scaling axes (data n, dimension d, groups G, MC samples) and compare dense inference against structured variational inference.

Where RE widens (paired)

Raise the source heterogeneity and watch the weakly identified quantities lift off the diagonal, the random-effects fit widening uncertainty selectively, not uniformly.

Leave-one-out across models

Cycle gaussian Fixed, gaussian RE, t RE and skew RE and watch the leave-one-out coverage jump from overconfident to well calibrated.

Figures 1 & 2: all fitted models versus the reference database

The paper fits six models (Gaussian, Student-t and skew-normal likelihoods, each in a fixed-effects and a hierarchical random-effects form) to synthetic glycolysis (Figure 1) and TCA-cycle (Figure 2) data, and compares them against the reference Database. All six recover the reference structure closely: the dominant between-model difference is in interval width, not in the posterior mean. Drag source heterogeneity and watch the random-effects fits widen uncertainty selectively (more for weakly identified quantities) rather than uniformly, exactly the source-level heterogeneity the hierarchical extension is built to capture.

Reading it. Each row is one kinetic quantity; the seven markers are the models in the legend (grey/open = fixed effects, coloured = random effects, blue = the Database reference). The markers cluster in the same place (the models agree on the central biochemical profile) but their bars differ in length (they disagree on uncertainty). Click a legend chip to show or hide a model. Push τ up: the random-effects bars (red, purple, green) stretch, and stretch more for the poorly constrained quantities, which is the hierarchical model quantifying source heterogeneity honestly while staying thermodynamically consistent. (Illustrative posteriors keyed to the real pathway parameters, to show the qualitative behaviour; the exact fitted values are in the paper.)

How random effects widen the interval: consensus versus per-source

Why do the random-effects fits carry wider intervals? A single kinetic constant is reported by several data sources that disagree. Fixed effects treat them as one population and pool to a tight interval; the hierarchical random-effects model adds a between-source variance τ2, so the pooled interval inflates to honestly cover the consensus. Raise the heterogeneity and watch the fixed-effects interval fail while the random-effects interval keeps its coverage.

Calibration and validation: does a 95% interval cover 95%?

A credible interval is only useful if it is honest: a 95% interval should contain the truth about 95% of the time. When the real measurement error is skewed or heavy-tailed, a Gaussian likelihood produces intervals that are too tight, so its empirical coverage falls below nominal, it is overconfident. The robust Student-t and skew-normal likelihoods, and the hierarchical random-effects pooling, keep the intervals honest. Drag the misspecification and watch the Gaussian curve fall away from the nominal line while the robust curves hold.

Left: empirical coverage of the chosen interval as the error departs from Gaussian (the paper's coverage-vs-skew and coverage-vs-df figures). Right: the reliability diagram, nominal against empirical coverage, where a well-calibrated model tracks the diagonal. Illustrative parametric curves keyed to the paper's qualitative results.

Figure 5: predictive coverage on real BRENDA data (leave-one-source-out)

This is the paper's flagship validation (Figure 5), reproduced from the actual fitted results. Every kinetic parameter in BRENDA is reported by many independent data sources, and leave-one-source-out (LOSO) cross-validation holds out one entire source, refits on the rest, predicts the held-out measurements, and checks how often the predictive interval at each nominal level actually covers them. A perfectly calibrated model tracks the dashed diagonal. The six models are the three likelihood families (Gaussian, skew-normal, Student-t), each as a fixed-effects fit (dashed line, open square) and a hierarchical random-effects fit (solid line, filled circle). The finding: fixed-effects models under-cover across the whole spectrum (worst in the upper tail), because they ignore between-source variance; adding source-level random effects lifts every family back toward nominal. Among the hierarchical fits the Gaussian and skew-normal models are nearly indistinguishable and closest to nominal, while Student-t improves markedly over its fixed-effects analogue and sits slightly below nominal at the highest levels. The calibration gain comes from modelling source heterogeneity, not from changing the mean prediction.

Random effects on

These curves are the exact fitted results from the paper's Julia pipeline (leave_one_source_out_cv on raw BRENDA glycolysis, then mean(pred_coverage) per model): 2,625 held-out predictions across 230 data sources, at 21 nominal levels. The fitted between-source standard deviation for the random-effects fits is τ̂ ≈ 1.29–1.37 (τ̂² = 1.66–1.89); the fixed-effects fits set it to zero. Colours match the paper: Gaussian steel-blue, skew-normal forest-green, Student-t firebrick.

Computational scaling: reaching genome-scale networks

Can this run on a whole metabolic network? Dense/exact inference cost grows steeply with the problem size, but structured variational inference exploits the sparse thermodynamic constraint graph and scales far more gently. Switch the horizontal axis to see the scaling in the number of data points, the model dimension, the number of reaction groups, or the Monte-Carlo sample budget.

Fixed versus random effects, paired

Does the hierarchical model widen every interval, or only the ones that need it? This paired view (paper Figure 7) plots each kinetic quantity's fixed-effects interval width against its random-effects width. Points sit on or above the diagonal, and the ones furthest above are the weakly identified quantities: the random-effects fit widens uncertainty exactly where source heterogeneity bites, not uniformly. Switch to RE vs RE (Figure 8) to see that two robust likelihoods agree on the uncertainty they assign.

Dot colour = identifiability: teal = well constrained (near the line), amber = weakly constrained (widened most). Illustrative widths keyed to the real pathway parameters.

Leave-one-out validation: hold out a record, predict it

The cleanest honesty check (the graphical-abstract validation, and the paper's leave-one-out panels): remove one measured record, refit, and predict the held-out value from the rest. A well-calibrated model should cover the held-out value about 95% of the time. A fixed-effects Gaussian ignores between-source variance, so its predictive interval is too tight and misses too often (red); a robust random-effects fit widens to stay honest. Raise the misspecification and watch the leave-one-out coverage.

What the demo shows

In PGI mode, four kinetic parameters for phosphoglucose isomerase (G6P ⇌ F6P), the forward and reverse turnover numbers kcat+, kcat and the Michaelis constants KM(G6P), KM(F6P), are measured directly, while the equilibrium constant Keq is determined from them by the Haldane relation. In Pathway mode a second reaction (F6P ⇌ FBP) is added, and the two equilibria share the formation energy of F6P, so measuring one reaction constrains the other.

1 · Balance

Consistency by construction

The posterior lives in the constrained subspace, so the balanced estimate satisfies every thermodynamic relation exactly, which the raw measurements never do.

2 · Quantify

Calibrated uncertainty

Every quantity gets a full posterior credible interval, and information flows through the constraints, so even unmeasured quantities are inferred.

3 · Robustify

Resist outliers

Heavy-tailed Student-t and asymmetric skew-normal likelihoods down-weight discordant measurements, keeping intervals honest under misspecification.

This is a lightweight illustration: exact conjugate linear-Gaussian inference, with Student-t and skew-normal handled by iterative reweighting (EM). It is not the full framework. For the model, its robust and hierarchical extensions, and the Julia package, see Bayesian Parameter Balancing Enables Robust and Consistent Estimation of Kinetic Parameter Uncertainty and the code on GitHub.