Applications

Methods meeting the real world

Where my mixture-of-experts theory, model-selection algorithms, simulation-based inference and efficient deep-learning methods leave the page and act on real data. Each entry names the actual dataset, the method applied, the scientific insight it produced, and the positive impact it is meant to have, then deep-links to the paper's own experiments section and lists all the related work.

My research develops principled statistical machine learning: guarantees for mixtures of experts, model selection without arbitrary sweeps, and honest uncertainty when the likelihood is intractable. The point of that rigour is to be trusted on problems that matter. The applications below span systems biology, crop science, medical AI, chemistry, planetary science and the sustainability of AI itself. Filter by area or search by dataset, method or keyword.

Systems and synthetic biology

Kinetic models are central to systems biology, but the enzyme-kinetic parameters compiled from the literature are incomplete, inconsistent and measured under heterogeneous conditions, so their uncertainty is easy to underestimate.

Robust Bayesian estimationBRENDA glycolysis & TCA-cycle kinetics

Bayesian parameter balancing enables robust and consistent estimation of kinetic parameter uncertainty

Method. A robust Bayesian reformulation of parameter balancing with heavy-tailed (Student-t / skew-normal) likelihoods and source-level random effects, fusing curated kinetic records (BRENDA, SABIO-RK) with thermodynamic constraints from eQuilibrator, with consistency guarantees.

Insight. On raw BRENDA glycolysis records (and the TCA cycle, with a community S. cerevisiae model), classical Gaussian credible intervals under-cover under heavy-tailed, skewed measurement noise; the robust likelihoods plus source random effects restore near-nominal 95% coverage, with most of the gain coming from modeling source heterogeneity, validated by leave-one-source-out predictive coverage.

Impact. Trustworthy kinetic models are the input to metabolic engineering, synthetic biology and drug-metabolism analysis; honest uncertainty keeps their predictions safe to act on.

Genomics, proteomics and crop resilience

High-dimensional biological measurements (thousands of proteins or genes per sample, few samples) are exactly the regime where the number and identity of the sub-populations, and which variables define them, is the scientific question.

Dendrograms of mixing measures233 maize genotypes, drought proteomics

Dendrograms of mixing measures for softmax-gated Gaussian mixture of experts

Method. Read the number of experts off a single over-fitted model by agglomerating the mixing measure, with finite-sample convergence rates and no model sweeps.

Insight. On a maize drought-response study (233 dent-maize genotypes under two watering conditions, seven ecophysiological traits), the method selects two experts and exposes an interpretable genotype-phenotype hierarchy, and stays robust under contamination where AIC, BIC and ICL over-select as the sample grows.

Impact. Principled, sweep-free order selection makes genotype-phenotype mapping for crop drought resilience reproducible and affordable, one of the levers for food security under a changing climate.

Variable selection under MNARArabidopsis thaliana transcriptome

Unified variable selection for model-based clustering with missing-not-at-random data

Method. A penalized Gaussian-mixture clustering with a data-driven penalty for variable selection and an MNARz mechanism that models the dependence between missingness and the latent class, with proven selection and asymptotic consistency.

Insight. On a real Arabidopsis thaliana transcriptome, the method partitions the data into 18 clusters and reaffirms a small set of core driver variables (with the rest explained by them), consistent with prior analyses, while running faster as the missing-data proportion rises.

Impact. Reliable subgroup and biomarker discovery from imperfect omics data, where the fact that a value is missing is itself informative.

Medical AI and healthcare

Medical foundation models must be reliable and data-efficient, because expert-annotated clinical data are scarce and mistakes are costly.

Graph-aligned vision-language modelVQA-RAD, SLAKE, PathVQA

ExGra-Med: extended context graph alignment for medical vision-language models

Method. A multi-graph alignment (images, instruction responses and extended captions) for a LLaMA-7B medical assistant, pre-trained across five imaging modalities from PMC-15M (chest X-ray, CT, MRI, histopathology and gross pathology).

Insight. On the medical visual-question-answering benchmarks VQA-RAD, SLAKE and PathVQA, ExGra-Med matches full-data performance using only 10% of pre-training data (a 20.13% gain on VQA-RAD), where LLaVA-Med drops from 72.6% to 52.4% under the same budget, and it surpasses strong baselines such as BioMedGPT and RadFM.

Impact. Data-efficient, better-grounded medical assistants lower the cost of building trustworthy clinical decision-support tools.

Chemistry and drug discovery

A molecule is not one shape but an ensemble of conformers; predicting its properties well means integrating that 3D structural information at scale.

E(3)-invariant conformer aggregationMoleculeNet + SARS-CoV screening

Structure-aware E(3)-invariant molecular conformer aggregation networks (CONAN-FGW)

Method. Aggregate a molecule's conformers through a differentiable fused Gromov-Wasserstein barycenter, giving an E(3)-invariant representation, paired with an efficient distance-geometry conformer generator.

Insight. On the MoleculeNet regression benchmarks Lipo, ESOL, FreeSolv and BACE (lipophilicity, aqueous solubility, hydration free energy, BACE-1 inhibition) and on SARS-CoV (Cov-2, Cov-2 3CL) classification, the method sets state-of-the-art property prediction while using significantly fewer 3D conformers.

Impact. Better property prediction and antiviral (SARS-CoV) screening accelerate the early, expensive stages of drug and materials discovery.

Fragment-aware conformer transformerMoleculeNet + MARCEL (Kraken, Drugs-75K)

FACET: a fragment-aware conformer ensemble transformer

Method. A structure-guided graph transformer that builds conformer ensembles from 2D graphs and reasons over molecular fragments, with 2D-3D conformer fusion supervised by fused Gromov-Wasserstein distances, replacing static geometric solvers.

Insight. On MoleculeNet (ESOL, FreeSolv, BACE) and the MARCEL benchmark (Kraken and the 52,569-molecule Drugs-75K), FACET reaches state-of-the-art accuracy while training 5-6 times cheaper than the conformer-aggregation baseline (214 versus 1,107 GPU-hours).

Impact. Cheap, accurate surrogates for quantum-mechanical conformer analysis make large-scale virtual screening feasible on a modest budget.

Planetary science and inverse problems

In many physical sciences the forward model is a simulator, not a likelihood: given parameters you can simulate data, but you cannot write down the density needed for standard inference.

Surrogate posteriors for ABCMars surface reflectance (Hapke model)

Summary statistics and discrepancy measures for approximate Bayesian computation via surrogate posteriors

Method. Use a fitted mixture-of-experts (GLLiM) surrogate posterior as a functional summary for approximate Bayesian computation, compared through L2 and optimal-transport (Wasserstein) distances, sidestepping hand-crafted summaries; the quasi-posterior is proven to converge to the true posterior.

Insight. The approach is especially effective when the true posterior is multimodal, where sample-based summaries struggle; it is validated on a planetary-science inversion (recovering the parameters of the Hapke reflectance model from Mars remote-sensing data) and on binaural sound-source localization.

Impact. Likelihood-free inference lets scientists read physical quantities off remote-sensing and acoustic data where no tractable likelihood exists.

Trustworthy and efficient AI

Large models are powerful but expensive; making them cheaper, greener and more reliable widens who can use them and lowers their environmental cost.

Spectrum-preserving token mergingImageNet, Flickr30k/MS-COCO, LLaVA

PiToMe: protecting token information with spectrum-preserving token merging

Method. An energy score flags large clusters of similar tokens for merging while preserving the low-energy (unique) tokens, provably preserving the informative spectrum.

Insight. Across ImageNet classification (ViT-MAE-H), Flickr30k and MS-COCO retrieval (CLIP/BLIP) and a LLaVA-7B assistant, PiToMe saves 40-60% of FLOPs with only a 0.5% accuracy drop on ImageNet (versus 2.6% for prior merging) and 0.3% on CLIP retrieval (versus 4.5%).

Impact. Cutting foundation-model inference compute in half lowers cost and energy, and puts capable models within reach of modest hardware.

Sparse mixture-of-experts routingenwik8, WikiText-103; SST, IMDB, BANKING77

HyperRouter: towards efficient training and inference of sparse mixture of experts

Method. Generate the router parameters from a fixed hypernetwork plus trainable embeddings, balancing trained against frozen routing to stabilise the notoriously brittle training of sparse mixture-of-experts language models.

Insight. On enwik8 and WikiText-103 language modeling and transfer to SST-2, SST-5, IMDB and BANKING77, HyperRouter reaches lower perplexity and better downstream accuracy than existing routing methods at the same number of inference experts.

Impact. Reliable sparse routing is how large language models scale capacity without scaling cost.

Token merging for SAMSegment Anything

StructSAM: structure- and spectrum-preserving token merging for Segment Anything models

Method. Structure- and spectrum-preserving token merging tailored to the Segment Anything foundation model, for resource-adaptive inference.

Insight. The merging keeps segmentation quality while reducing the token budget, so the same model runs under a range of compute constraints.

Impact. Adaptive segmentation lets a single foundation model serve everything from a data centre to an edge device.

Competition-based routingSparse MoE benchmarks

CompeteSMoE: statistical guarantees and efficient architecture for sparse mixtures of experts

Method. Let experts compete for tokens, with statistical guarantees for the resulting routing, an alternative to the usual learned gate.

Insight. Competition routing improves the effective use of experts in sparse mixture-of-experts models, backed by convergence theory.

Impact. Principled routing makes the efficiency gains of sparse experts dependable rather than a tuning lottery.

The theory behind the applications. These results rest on a programme of mathematical guarantees: universal approximation of conditional densities by mixtures of experts (softmax gating, Lebesgue-space approximation), convergence rates for parameter estimation (Gaussian-gated, logistic-gated), non-asymptotic model selection (slope heuristic, joint rank and variable selection, risk bounds), and the large-sample behaviour of Bayesian model evaluation (DIC and WBIC). The full list is on the Publications page, and the methods are runnable in the Software section.