Optimizing yield in pharmaceutical fermentation presents unique challenges for data-driven modelling. Batch cycles span several days, with process parameters changing continuously in real time, while final API yield is measured only at the end of the batch. This delayed feedback prevents in-process correction and limits the effectiveness of traditional real-time control methods. Complex interactions among variables, combined with limited historical batch data, further create a high-noise, data-scarce environment that constrains conventional modelling approaches.

This paper presents a structured machine learning approach, built in JMP, to investigate the causes of fermentation yield decline and to identify strategies for yield recovery and optimization in a pharmaceutical bioprocess. To address the delayed-outcome problem, the fermentation process is divided into discrete temporal stages, enabling stage-wise analysis of process behaviour. This staged structure links parameter performance at each stage to the final yield outcome, creating surrogate indicators of batch health that can be monitored before the batch concludes.

Historical high-yield batches are benchmarked against underperforming batches using descriptive statistics, distribution analysis, and process capability studies to identify critical deviations. Key parameters, including feed rates, fatty acid concentrations, pH, dissolved oxygen, dosage variables, and raw material consistency, are analysed individually and in combination. Quartile thresholds and inhibitory limits are established to define actionable control ranges for each parameter.

Three machine learning models, linear regression, CART, and random forest, are applied within an ensemble framework in JMP, followed by experimentation using a Bayesian optimization approach. Together, these models capture linear trends, interpretable decision rules, and nonlinear interactions to identify optimal operating ranges, replacing trial-and-error experimentation with a data-driven, systematic approach to process optimization.

The stage-wise ensemble approach successfully separated the process signals driving yield decline from normal batch-to-batch noise, and identified clear, actionable operating ranges, quartile thresholds and inhibitory limits, for feed rates, fatty acid concentrations, pH, dissolved oxygen, dosage variables, and raw material consistency at each fermentation stage. Batches operated within these stage-specific ranges showed measurably higher and more consistent activity than historical underperforming batches, and the Bayesian optimization step further refined the recommended operating points beyond what the tree-based and regression models alone identified, converging on near-optimal conditions without exhaustive additional experimentation.

A central goal of this work was democratization: turning a modelling exercise that would typically require a dedicated data science team into a workflow that process engineers and operations personnel can run and interpret themselves. By packaging the staged benchmarking, ensemble modelling, and Bayesian optimization entirely within JMP's interactive platform, the analysis and its outputs, control ranges, stage-wise recommendations, and a Profiler-based digital twin, are directly accessible to the people who run the fermentation process day to day. This shifts yield optimization from a periodic, specialist-led project to a continuously available, self-service capability, and offers a template that other JMP users can apply to their own delayed-outcome, high-variability processes.

0 Comments
Presented At Discovery Summit 2026

Presenter

Skill level

Intermediate
  • Beginner
  • Intermediate
  • Advanced

Skill level

Advanced
  • Beginner
  • Intermediate
  • Advanced
Published on ‎07-16-2026 11:13 AM by Community Manager Community Manager | Updated on ‎08-28-2026 06:49 AM

Optimizing yield in pharmaceutical fermentation presents unique challenges for data-driven modelling. Batch cycles span several days, with process parameters changing continuously in real time, while final API yield is measured only at the end of the batch. This delayed feedback prevents in-process correction and limits the effectiveness of traditional real-time control methods. Complex interactions among variables, combined with limited historical batch data, further create a high-noise, data-scarce environment that constrains conventional modelling approaches.

This paper presents a structured machine learning approach, built in JMP, to investigate the causes of fermentation yield decline and to identify strategies for yield recovery and optimization in a pharmaceutical bioprocess. To address the delayed-outcome problem, the fermentation process is divided into discrete temporal stages, enabling stage-wise analysis of process behaviour. This staged structure links parameter performance at each stage to the final yield outcome, creating surrogate indicators of batch health that can be monitored before the batch concludes.

Historical high-yield batches are benchmarked against underperforming batches using descriptive statistics, distribution analysis, and process capability studies to identify critical deviations. Key parameters, including feed rates, fatty acid concentrations, pH, dissolved oxygen, dosage variables, and raw material consistency, are analysed individually and in combination. Quartile thresholds and inhibitory limits are established to define actionable control ranges for each parameter.

Three machine learning models, linear regression, CART, and random forest, are applied within an ensemble framework in JMP, followed by experimentation using a Bayesian optimization approach. Together, these models capture linear trends, interpretable decision rules, and nonlinear interactions to identify optimal operating ranges, replacing trial-and-error experimentation with a data-driven, systematic approach to process optimization.

The stage-wise ensemble approach successfully separated the process signals driving yield decline from normal batch-to-batch noise, and identified clear, actionable operating ranges, quartile thresholds and inhibitory limits, for feed rates, fatty acid concentrations, pH, dissolved oxygen, dosage variables, and raw material consistency at each fermentation stage. Batches operated within these stage-specific ranges showed measurably higher and more consistent activity than historical underperforming batches, and the Bayesian optimization step further refined the recommended operating points beyond what the tree-based and regression models alone identified, converging on near-optimal conditions without exhaustive additional experimentation.

A central goal of this work was democratization: turning a modelling exercise that would typically require a dedicated data science team into a workflow that process engineers and operations personnel can run and interpret themselves. By packaging the staged benchmarking, ensemble modelling, and Bayesian optimization entirely within JMP's interactive platform, the analysis and its outputs, control ranges, stage-wise recommendations, and a Profiler-based digital twin, are directly accessible to the people who run the fermentation process day to day. This shifts yield optimization from a periodic, specialist-led project to a continuously available, self-service capability, and offers a template that other JMP users can apply to their own delayed-outcome, high-variability processes.



0 Kudos