Hi @FrancesBudgie97,
Welcome in the Community !
It's quite rare to see Taguchi designs used. Just to be sure, your three factors are numerical continuous, but to create the design, you have defined them as categorical or discrete numeric ?
I wouldn't remove by default any terms that may be correlated to others in the model. Correlation between terms are often found in designs, and that doesn't prevent them to be included in the model. The consequence will be some restrictions in the terms inclusion (you won't be able to include them all, or you'll end up with a singularity in your model), some lack of precision to estimate them (with broader confidence intervals around the estimate values) and some level of collinearity (you can check Variance Inflation Factors (VIF) values for the terms included in your model to assess the degree of collinearity in your model).
Since you can't estimate all possible terms of a Response Surface Model (RSM) from your design, you are in a supersaturated situation. You'll need to use some specific methods that help determine which terms may be the most impactful ones on your different responses. Some interesting options are:
- The Fit Two Level Screening Platform: Despite its name, I often used it for situations where I have three levels for my factors, and not enough runs to estimate a full RSM. The platform relies on simulations to determine which terms may have the biggest influence on the response. To do this, the platform follows some rules (Statistical Details for Order of Effect Entry), like the Effect Hierarchy and Effect Heredity principles.
- Generalized Regression Models (with JMP Pro): The Generalized Regression models in JMP Pro have various Estimation Method Options (like Pruned Forward Selection, Two Stage Forward Selection, or Lasso/Elastic Net/...) that can help screen important effects from a large list of effects.
- Stepwise Regression Models (with JMP): I wouldn't recommend relying too much on stepwise model, since they will "brute-force" a way to find the best model (sometimes overfitted) based on many different terms combinations. However, they may be useful as a model comparison tool, particularly if you can create multiple models to create Raster plots, in order to compare which terms are often included in many different models.
More info about the Raster plots: Linkedin Post
Recording Experimenters' Club Q2 2026_Beyond One Best Model: What a DSD Can Really Tell You
You can find one use case presented in the french JMP User group here, where I used these different platforms to compare model fit metrics (R² / RMSE / ...), model complexity/accuracy balance (with information criterion like Likelihood, AICc, and BIC), agreement between models for terms inclusions, etc...
Concerning your question about how to add/remove terms to build your models, the methods mentioned above can help. But more importantly, you need to define :
- Your objective: Are you in a screening situation (where p-values may help to differentiate true signal from random variation or noise) ? Or in an optimization/prediction scenario (where accuracy of the model may matter more, so you may rely on RMSE and other predictive metrics)? Or you are in the middle, and not sure (Information criterion may help you compare models and determine the right balance for the terms and number of terms included by balancing accuracy with complexity).
- Your domain knowledge/expertise: Are you able to remove/add some terms based on historical data or knowledge about the system you're studying ? Are there already some interactions or main effects known to be active and important ? Can you check/assess a model with domain experts to see if the model makes some sense ?
- Your validation strategy: Once you have selected models that seem to be adequate for your responses, you need to validate these models. Based on your objective(s) and experimental budget, you may have different scenarii to make validation runs. Are you only interested in validating the optimum points ? Or do you want to validate the model on the factors ranges studied (and to do so, run combination of factors levels not in the design but respecting the factors ranges) ? There are no strict rules to validate your findings, but taking extra care (and time !) to run these validation experiments can greatly strenghten your confidence in the models (and the confidence from your reviewers !).
Hope this answer will help you,
Victor GUILLER
"It is not unusual for a well-designed experiment to analyze itself" (Box, Hunter and Hunter)