cancel
Showing results for 
Show  only  | Search instead for 
Did you mean: 
  • Don’t miss your chance to experience Discovery Summit Europe at our best available rate! Early bird registration through 31 Oct.
  • Graph Builder Toolbar streamlines interactive graphing. Creates shortcuts on Report and GraphBuilder Helpers toolbars. Download and install the extension.
  • JMP will suspend normal business operations for our Rest and Recharge Day on Friday, October 2, 2026.
    Regular business hours will resume on Monday, October 5, 2026.

Discussions

Solve problems, and share tips and tricks with other JMP users.
Choose Language Hide Translation Bar

How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?)

Hi there,

I am trying to find the best way to select the terms to include in my quadratic model.  :)

First off, I am using Least Squares Method to understand how pH and other continuous (also maybe some discrete indicators to add to my model soon) are affected by three different factors (concentration A , concentration B and ratio C). So I need to find the relationship between the factors (their square and their interactions) and the outputs that I have measured (pH, rheology values etc.)

I got my data by doing an L9 Taguchi matrix design of experiments (three parameters that have three levels each) so I only have 9 experiments. I have to choose the terms for my models used for each output. (A first model for pH, another can be used for rheology measurments etc.)

With only 9 runs, I know I cannot fit the full quadratic model (10 parameters). The main effects are orthogonal, but the two-factor interactions are aliased with main effects (|r| ≈ 0.58), so I plan to drop them. That leaves main effects + quadratic terms (6 parameters). My question is really about how to select among these remaining terms: manual reduction (removing non-significant terms one by one) vs LASSO — which is more defensible for a PhD?

Thanks for any help!

Kind regards,

Anna

2 REPLIES 2

Re: How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?)

Feel free to ask for any extra details !

Victor_G
Super User

Re: How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?)

Hi @FrancesBudgie97,

Welcome in the Community !

It's quite rare to see Taguchi designs used. Just to be sure, your three factors are numerical continuous, but to create the design, you have defined them as categorical or discrete numeric ?

I wouldn't remove by default any terms that may be correlated to others in the model. Correlation between terms are often found in designs, and that doesn't prevent them to be included in the model. The consequence will be some restrictions in the terms inclusion (you won't be able to include them all, or you'll end up with a singularity in your model), some lack of precision to estimate them (with broader confidence intervals around the estimate values) and some level of collinearity (you can check Variance Inflation Factors (VIF) values for the terms included in your model to assess the degree of collinearity in your model).

Since you can't estimate all possible terms of a Response Surface Model (RSM) from your design, you are in a supersaturated situation. You'll need to use some specific methods that help determine which terms may be the most impactful ones on your different responses. Some interesting options are:

You can find one use case presented in the french JMP User group here, where I used these different platforms to compare model fit metrics (R² / RMSE / ...), model complexity/accuracy balance (with information criterion like Likelihood, AICc, and BIC), agreement between models for terms inclusions, etc...

Concerning your question about how to add/remove terms to build your models, the methods mentioned above can help. But more importantly, you need to define :

  • Your objective: Are you in a screening situation (where p-values may help to differentiate true signal from random variation or noise) ? Or in an optimization/prediction scenario (where accuracy of the model may matter more, so you may rely on RMSE and other predictive metrics)? Or you are in the middle, and not sure (Information criterion may help you compare models and determine the right balance for the terms and number of terms included by balancing accuracy with complexity).
  • Your domain knowledge/expertise: Are you able to remove/add some terms based on historical data or knowledge about the system you're studying ? Are there already some interactions or main effects known to be active and important ? Can you check/assess a model with domain experts to see if the model makes some sense ?
  • Your validation strategy: Once you have selected models that seem to be adequate for your responses, you need to validate these models. Based on your objective(s) and experimental budget, you may have different scenarii to make validation runs. Are you only interested in validating the optimum points ? Or do you want to validate the model on the factors ranges studied (and to do so, run combination of factors levels not in the design but respecting the factors ranges) ? There are no strict rules to validate your findings, but taking extra care (and time !) to run these validation experiments can greatly strenghten your confidence in the models (and the confidence from your reviewers !).

Hope this answer will help you,

Victor GUILLER

"It is not unusual for a well-designed experiment to analyze itself" (Box, Hunter and Hunter)

Recommended Articles