Research

How does a system change its behavior?

My current research centers on biological systems: enzyme function, metabolism, and cellular adaptation. The methodological thread is the integration of learning, mechanistic constraints, and experimental evidence.

The Persistent Gap: Data-rich models can detect patterns, while mechanistic models can enforce feasibility. Yet neither alone reliably explains or redesigns complex biological behavior. My research portfolio addresses that gap across sectors and scales.

The Consistent Scientific Objective: Determine why a biological state is possible, which constraints sustain it, and how those constraints can be measured, challenged, or redesigned.

Research Themes

Enzymes & Metabolism

ML integrates sequence and substrate chemistry to estimate kinetic regimes and mutation-sensitive catalysis. Isotope-resolved data and perturbation modeling recover transient pathway activity and identify control points governing homeostasis.

Disease & Constraints

Multi-omics data are embedded within stoichiometric, thermodynamic, and enzyme-capacity constraints. This distinguishes statistical signatures from metabolically feasible states, exposing context-specific vulnerabilities in disease networks.

Plants & Synthetic Biology

Proteome-aware models reveal how nutrient availability and redox balance shape adaptation. These insights guide target selection, pathway redesign, and the development of highly robust bioproduction hosts.

Emerging Frontiers

Expanding this body of work through molecular simulation, agentic scientific workflows, and LLM-assisted evidence synthesis.

Case Studies

Enzyme Engineering (Kinetics) / Scientific AI

When a mutation changes the enzyme, the prediction should change too.

Mutation-sensitive predictions of enzyme kinetic regimes, built around protein sequence, substrate chemistry, and curated evidence. CatRange (PNAS Nexus, 2026) and its precursor framework RealKcat (bioRxiv, 2025) together form this line of work.

Published + public software Question → model → evidence
PROTEIN SEQUENCE
→ Variant
SUBSTRATE CHEMISTRY
→ Context
k₁ k₂ k₃ k₄ k₅

Regimes, not false precision.

Conceptual Schematic · Not Experimental Output

The Question

An enzyme can look almost identical in sequence and behave very differently after a catalytic-site mutation. How can a model reflect that difference without implying more numerical precision than the evidence supports?

My Contribution

I developed CatRange with collaborators, connecting biochemical data curation, enzyme–substrate representations, model development, evaluation, and usable inference.

  • Protein representations
  • Substrate chemistry
  • Gradient-boosted classification
  • Mutation-aware evaluation
  • Research inference

What the record shows

The PNAS Nexus publication and public repository provide the research and software record. The contribution is a range-based, mutation-sensitive modeling framework—not an experimental measurement of every predicted variant.

What this does—and does not—establish

A predicted regime is not a measured kinetic constant. Neighboring-bin recovery is an evaluation criterion, not automatically a calibrated confidence interval. Assay context, sequence coverage, and the deployed bin definitions still matter.

Biological Systems / Proteomics

A changing proteome. A more persistent growth-associated core.

Separating shared growth-associated proteins from environment-specific adaptation in a metabolically versatile bacterium.

Published + public software Question → model → evidence
Growth-associated core
01 · Cross-condition core
02 · Adaptive regulators
03 · Substrate specialists
04 · Conditional hubs

Conceptual Schematic · Not Experimental Output

The Question

Rhodopseudomonas palustris reorganizes its proteome across lignin-derived substrates and oxygen conditions. Which patterns remain informative about growth across those different environments?

My Contribution

I developed CorePredX to connect quantitative proteomics with growth prediction and dependence-aware interpretation. My work spans the computational question, methodology, software, analysis, validation, and scientific communication.

  • Quantitative proteomics
  • Neural prediction
  • SHAP attribution
  • Redundancy analysis
  • Cross-condition interpretation

What the record shows

The mSystems study identifies a compact hierarchy of candidate growth-associated proteins beneath broader proteome remodeling, motivating more focused biological follow-up.

What this does—and does not—establish

Predictive importance does not establish a causal regulator or a directed regulatory edge. Correlated features, the sampled environments, and independent experimental validation constrain the interpretation.

Metabolic Engineering / Dynamic Modeling

Learning what a pathway is doing when the measurements are imperfect.

Isotope-informed dynamic modeling of plant sphingolipid metabolism, with uncertainty treated as part of the scientific problem.

Published Question → model → evidence

The Question

How do you quantify transient metabolic fluxes when the experimental data is sparse and noisy?

My Contribution

r-DMFA framework, 15N isotope-labeling, targeted metabolomics, dynamic flux sampling. Identified SBH, LCBK control nodes.

  • Isotope labeling integration
  • Dynamic flux analysis
  • Uncertainty quantification
  • Control point identification
  • Arabidopsis metabolism

What the record shows

iScience publication demonstrates flux identification under data variability.

What this does—and does not—establish

Point estimates of fluxes carry uncertainty; validated against biological priors, not independent kinetic measurements.

Systems Immunology / Multi-Omics

Connecting early immune signals to later antibody trajectories.

DynPertBoost: temporal, multi-horizon ensemble framework for pertussis vaccination response.

Research software Question → model → evidence

The Question

Can early cytokine and gene-expression patterns predict vaccine-induced antibody responses at later timepoints?

My Contribution

Built multi-horizon ensemble integrating cytokine, transcriptomic, and antibody data.

What the record shows

Demonstrates predictive utility of early-stage multi-omics integration in vaccine response forecasting.

What this does—and does not—establish

Identifies predictive biomarkers but does not directly prove mechanistic causality without further biological validation.

Cancer Biology / Metabolic Modeling

How do cancer cells adapt when several resources become scarce?

Time-resolved analysis of pancreatic-cancer adaptation to hypoxia, acidity, and nutrient limitation.

In progress Question → model → evidence

The Question

What metabolic pathways do pancreatic cancer cells activate when simultaneously stressed by low oxygen, acidic pH, and nutrient deprivation?

My Contribution

Building temporally resolved metabolic models, analyzing PDAC RNA-seq.

What the record shows

Reveals shifting dependencies across multiple simultaneous stress axes, informing context-specific vulnerabilities.

What this does—and does not—establish

Models suggest potential therapeutic targets based on transcriptomic shifts, which require subsequent empirical targeting to confirm essentiality.

Resources

Manuscript in preparation.

Process Systems Engineering

Making physical processes do what you designed them to do.

Nonlinear model predictive control and ANN-based controllers on instrumented process platforms.

Published Question → model → evidence

The Question

How do you design a controller that handles nonlinear behavior, model uncertainty, and real-time constraints simultaneously?

My Contribution

Designed/fabricated cascaded three-tank system, implemented NMPC, MHE, EKF, economic MPC. Evaluated ANN-based predictive control.

What the record shows

Demonstrated real-time constraint satisfaction and setpoint tracking in physical hardware under disturbances.

What this does—and does not—establish

Validates controller architectures on specific pilot-scale hardware, not universal bounds on stability under arbitrary noise.

Internet-of-Things (IoT) / Smart Agriculture

Measure the variation that matters—not every location equally.

Machine-learning-based sensor clustering for controlled-environment agriculture.

Published Question → model → evidence

The Question

If a greenhouse has 56 sensor locations, how do you find the minimum set that captures meaningful variation without redundant monitoring?

My Contribution

Online K-Means++ clustering, psychrometric feature construction, multi-season validation.

What the record shows

Demonstrates reliable microclimate estimation from a reduced sensor footprint, lowering deployment costs.

What this does—and does not—establish

Optimizes sensor placement for specific geometries and seasons; re-calibration may be necessary for structural changes.

Scientific AI / Agents

Make the evidence travel with the recommendation.

Evidence-linked mutation analysis that brings scientific language models, biochemical context, and tools into one research workflow.

In development Question → model → evidence

The Question

How can an AI assistant help a researcher find, evaluate, and act on biochemical evidence about enzyme mutations?

My Contribution

Training/fine-tuning LLMs with biochemical evidence. Building research prototype.

What the record shows

Early integration of biochemical heuristics and literature context inside an accessible chat agent.

What this does—and does not—establish

Provides augmented literature exploration, not a substitute for human scientific judgment or rigorous simulation.

Resources

Prototype in development.

Start a conversation. Let’s work on a difficult problem.

Research collaboration, scientific software, engineering questions—or a thoughtful conversation about where to begin.

Get in touch