<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?) in Discussions</title>
    <link>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972590#M110572</link>
    <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.jmp.com/t5/user/viewprofilepage/user-id/114120"&gt;@FrancesBudgie97&lt;/a&gt;,&lt;/P&gt;
&lt;P&gt;Welcome in the Community !&lt;/P&gt;
&lt;P&gt;It's quite rare to see Taguchi designs used. Just to be sure, your three factors are numerical continuous, but to create the design, you have defined them as categorical or discrete numeric ?&lt;/P&gt;
&lt;P&gt;I wouldn't remove by default any terms that may be correlated to others in the model. Correlation between terms are often found in designs, and that doesn't prevent them to be included in the model. The consequence will be some restrictions in the terms inclusion (you won't be able to include them all, or you'll end up with a &lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/special-reports.shtml#" target="_self"&gt;singularity&lt;/A&gt; in your model), some lack of precision to estimate them (with broader confidence intervals around the estimate values) and some level of collinearity (you can check &lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/parameter-estimates-for-original-predictors.shtml" target="_self"&gt;Variance Inflation Factors (VIF) values&lt;/A&gt; for the terms included in your model to assess the degree of collinearity in your model).&lt;/P&gt;
&lt;P&gt;Since you can't estimate all possible terms of a Response Surface Model (RSM) from your design, you are in a supersaturated situation. You'll need to use some specific methods that help determine which terms may be the most impactful ones on your different responses. Some interesting options are:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/the-fit-two-level-screening-platform.shtml" target="_blank"&gt;The Fit Two Level Screening Platform&lt;/A&gt;: Despite its name, I often used it for situations where I have three levels for my factors, and not enough runs to estimate a full RSM. The platform relies on simulations to determine which terms may have the biggest influence on the response. To do this, the platform follows some rules (&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/statistical-details-for-order-of-effect-entry.shtml#ww155172" target="_blank"&gt;Statistical Details for Order of Effect Entry&lt;/A&gt;), like the&amp;nbsp;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/effect-hierarchy.shtml#ww347026" target="_blank"&gt;Effect Hierarchy&lt;/A&gt;&amp;nbsp;and&amp;nbsp;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/effect-heredity.shtml#ww347029" target="_blank"&gt;Effect Heredity&amp;nbsp;&lt;/A&gt;principles.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/generalized-regression-models.shtml" target="_blank"&gt;Generalized Regression Models&lt;/A&gt;&amp;nbsp;(with JMP Pro): The Generalized Regression models in JMP Pro have various&amp;nbsp;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/estimation-method-options.shtml?_gl=1*1itgp8m*_up*MQ..*_ga*MTE1OTgyMzQwMi4xNzg5OTc2Nzgy*_ga_BRNVBEC1RS*czE3ODk5NzY3ODEkbzEkZzAkdDE3ODk5NzY3ODEkajYwJGwwJGgw#" target="_blank"&gt;Estimation Method Options&lt;/A&gt;&amp;nbsp;(like&amp;nbsp;&lt;SPAN&gt;Pruned Forward Selection,&amp;nbsp;Two Stage Forward Selection, or Lasso/Elastic Net/...)&lt;/SPAN&gt;&amp;nbsp;that can help screen important effects from a large list of effects.&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/stepwise-regression-models.shtml#" target="_blank"&gt;Stepwise Regression Models&lt;/A&gt;&amp;nbsp;(with JMP): I wouldn't recommend relying too much on stepwise model, since they will "brute-force" a way to find the best model (sometimes overfitted) based on many different terms combinations. However, they may be useful as a model comparison tool, particularly if you can create multiple models to create Raster plots, in order to compare which terms are often included in many different models.&lt;BR /&gt;More info about the Raster plots: &lt;A href="https://www.linkedin.com/posts/victorguiller_experimentersclub-designofexperiments-activity-7454786628448841732-ZFcw?utm_source=share&amp;amp;utm_medium=member_desktop&amp;amp;rcm=ACoAAA79OucB6Jgj4QVIgAHP5Ju6rirGp8XmPcI" target="_self"&gt;Linkedin Post&lt;/A&gt;&lt;BR /&gt;&lt;A href="https://community.jmp.com/t5/Design-of-Experiments-Club/Recording-Experimenters-Club-Q2-2026-Beyond-One-Best-Model-What/m-p/945035" target="_self"&gt;&lt;SPAN&gt;Recording Experimenters' Club Q2 2026_Beyond One Best Model: What a DSD Can Really Tell You&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;You can find one use case presented in the french JMP User group &lt;A href="https://community.jmp.com/t5/Groupe-francophone-des/D%C3%A9couverte-des-plans-OML-Orthogonal-Main-Effects-Screening/m-p/843161" target="_self"&gt;here&lt;/A&gt;, where I used these different platforms to compare model fit metrics (R² / RMSE / ...), model complexity/accuracy balance (with information criterion like &lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/likelihood-aicc-and-bic.shtml#ww293087" target="_blank"&gt;Likelihood, AICc, and BIC)&lt;/A&gt;, agreement between models for terms inclusions, etc...&lt;/P&gt;
&lt;P&gt;Concerning your question about how to add/remove terms to build your models, the methods mentioned above can help. But more importantly, you need to define :&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Your objective&lt;/STRONG&gt;: Are you in a screening situation (where p-values may help to differentiate true signal from random variation or noise) ? Or in an optimization/prediction scenario (where accuracy of the model may matter more, so you may rely on RMSE and other predictive metrics)? Or you are in the middle, and not sure (Information criterion may help you compare models and determine the right balance for the terms and number of terms included by balancing accuracy with complexity).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Your domain knowledge/expertise&lt;/STRONG&gt;: Are you able to remove/add some terms based on historical data or knowledge about the system you're studying ? Are there already some interactions or main effects known to be active and important ? Can you check/assess a model with domain experts to see if the model makes some sense ?&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Your validation strategy&lt;/STRONG&gt;: Once you have selected models that seem to be adequate for your responses, you need to validate these models. Based on your objective(s) and experimental budget, you may have different scenarii to make validation runs. Are you only interested in validating the optimum points ? Or do you want to validate the model on the factors ranges studied (and to do so, run combination of factors levels not in the design but respecting the factors ranges) ? There are no strict rules to validate your findings, but taking extra care (and time !) to run these validation experiments can greatly strenghten your confidence in the models (and the confidence from your reviewers !).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Hope this answer will help you,&lt;/P&gt;</description>
    <pubDate>Mon, 21 Sep 2026 08:18:40 GMT</pubDate>
    <dc:creator>Victor_G</dc:creator>
    <dc:date>2026-09-21T08:18:40Z</dc:date>
    <item>
      <title>How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?)</title>
      <link>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972266#M110563</link>
      <description>&lt;P data-pm-slice="1 1 []"&gt;Hi there,&lt;/P&gt;
&lt;P&gt;I am trying to find the best way to select the terms to include in my quadratic model.&amp;nbsp; :)&lt;/img&gt;&lt;/P&gt;
&lt;P&gt;First off, I am using&amp;nbsp;Least Squares Method to understand how pH and other continuous (also maybe some discrete indicators to add to my model soon) are affected by three different factors (concentration A ,&amp;nbsp;concentration B and ratio C). So I need to find the relationship between the factors (their square and their interactions) and the outputs that I have measured (pH, rheology values etc.)&lt;/P&gt;
&lt;P&gt;I got my data by doing an&amp;nbsp;L9 Taguchi matrix design of experiments (three parameters that have three levels each) so I only have 9 experiments.&amp;nbsp;I have to choose the terms for my models used for each output. (A first model for pH, another can be used for rheology measurments etc.)&lt;/P&gt;
&lt;P&gt;With only 9 runs, I know I cannot fit the full quadratic model (10 parameters). The main effects are orthogonal, but the two-factor interactions are aliased with main effects (|r| ≈ 0.58), so I plan to drop them. That leaves main effects + quadratic terms (6 parameters). My question is really about how to select among these remaining terms: manual reduction (removing non-significant terms one by one) vs LASSO — which is more defensible for a PhD?&lt;/P&gt;
&lt;P&gt;Thanks for any help!&lt;/P&gt;
&lt;P&gt;Kind regards,&lt;/P&gt;
&lt;P&gt;Anna&lt;/P&gt;</description>
      <pubDate>Fri, 18 Sep 2026 09:09:05 GMT</pubDate>
      <guid>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972266#M110563</guid>
      <dc:creator>FrancesBudgie97</dc:creator>
      <dc:date>2026-09-18T09:09:05Z</dc:date>
    </item>
    <item>
      <title>Re: How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?)</title>
      <link>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972275#M110564</link>
      <description>&lt;P&gt;Feel free to ask for any extra details !&lt;/P&gt;</description>
      <pubDate>Fri, 18 Sep 2026 09:09:59 GMT</pubDate>
      <guid>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972275#M110564</guid>
      <dc:creator>FrancesBudgie97</dc:creator>
      <dc:date>2026-09-18T09:09:59Z</dc:date>
    </item>
    <item>
      <title>Re: How to select model terms for small sample size (SHould I use Lasso? chose terms manually with p-value?)</title>
      <link>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972590#M110572</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.jmp.com/t5/user/viewprofilepage/user-id/114120"&gt;@FrancesBudgie97&lt;/a&gt;,&lt;/P&gt;
&lt;P&gt;Welcome in the Community !&lt;/P&gt;
&lt;P&gt;It's quite rare to see Taguchi designs used. Just to be sure, your three factors are numerical continuous, but to create the design, you have defined them as categorical or discrete numeric ?&lt;/P&gt;
&lt;P&gt;I wouldn't remove by default any terms that may be correlated to others in the model. Correlation between terms are often found in designs, and that doesn't prevent them to be included in the model. The consequence will be some restrictions in the terms inclusion (you won't be able to include them all, or you'll end up with a &lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/special-reports.shtml#" target="_self"&gt;singularity&lt;/A&gt; in your model), some lack of precision to estimate them (with broader confidence intervals around the estimate values) and some level of collinearity (you can check &lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/parameter-estimates-for-original-predictors.shtml" target="_self"&gt;Variance Inflation Factors (VIF) values&lt;/A&gt; for the terms included in your model to assess the degree of collinearity in your model).&lt;/P&gt;
&lt;P&gt;Since you can't estimate all possible terms of a Response Surface Model (RSM) from your design, you are in a supersaturated situation. You'll need to use some specific methods that help determine which terms may be the most impactful ones on your different responses. Some interesting options are:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/the-fit-two-level-screening-platform.shtml" target="_blank"&gt;The Fit Two Level Screening Platform&lt;/A&gt;: Despite its name, I often used it for situations where I have three levels for my factors, and not enough runs to estimate a full RSM. The platform relies on simulations to determine which terms may have the biggest influence on the response. To do this, the platform follows some rules (&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/statistical-details-for-order-of-effect-entry.shtml#ww155172" target="_blank"&gt;Statistical Details for Order of Effect Entry&lt;/A&gt;), like the&amp;nbsp;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/effect-hierarchy.shtml#ww347026" target="_blank"&gt;Effect Hierarchy&lt;/A&gt;&amp;nbsp;and&amp;nbsp;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/effect-heredity.shtml#ww347029" target="_blank"&gt;Effect Heredity&amp;nbsp;&lt;/A&gt;principles.&amp;nbsp;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/generalized-regression-models.shtml" target="_blank"&gt;Generalized Regression Models&lt;/A&gt;&amp;nbsp;(with JMP Pro): The Generalized Regression models in JMP Pro have various&amp;nbsp;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/estimation-method-options.shtml?_gl=1*1itgp8m*_up*MQ..*_ga*MTE1OTgyMzQwMi4xNzg5OTc2Nzgy*_ga_BRNVBEC1RS*czE3ODk5NzY3ODEkbzEkZzAkdDE3ODk5NzY3ODEkajYwJGwwJGgw#" target="_blank"&gt;Estimation Method Options&lt;/A&gt;&amp;nbsp;(like&amp;nbsp;&lt;SPAN&gt;Pruned Forward Selection,&amp;nbsp;Two Stage Forward Selection, or Lasso/Elastic Net/...)&lt;/SPAN&gt;&amp;nbsp;that can help screen important effects from a large list of effects.&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/stepwise-regression-models.shtml#" target="_blank"&gt;Stepwise Regression Models&lt;/A&gt;&amp;nbsp;(with JMP): I wouldn't recommend relying too much on stepwise model, since they will "brute-force" a way to find the best model (sometimes overfitted) based on many different terms combinations. However, they may be useful as a model comparison tool, particularly if you can create multiple models to create Raster plots, in order to compare which terms are often included in many different models.&lt;BR /&gt;More info about the Raster plots: &lt;A href="https://www.linkedin.com/posts/victorguiller_experimentersclub-designofexperiments-activity-7454786628448841732-ZFcw?utm_source=share&amp;amp;utm_medium=member_desktop&amp;amp;rcm=ACoAAA79OucB6Jgj4QVIgAHP5Ju6rirGp8XmPcI" target="_self"&gt;Linkedin Post&lt;/A&gt;&lt;BR /&gt;&lt;A href="https://community.jmp.com/t5/Design-of-Experiments-Club/Recording-Experimenters-Club-Q2-2026-Beyond-One-Best-Model-What/m-p/945035" target="_self"&gt;&lt;SPAN&gt;Recording Experimenters' Club Q2 2026_Beyond One Best Model: What a DSD Can Really Tell You&lt;/SPAN&gt;&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;You can find one use case presented in the french JMP User group &lt;A href="https://community.jmp.com/t5/Groupe-francophone-des/D%C3%A9couverte-des-plans-OML-Orthogonal-Main-Effects-Screening/m-p/843161" target="_self"&gt;here&lt;/A&gt;, where I used these different platforms to compare model fit metrics (R² / RMSE / ...), model complexity/accuracy balance (with information criterion like &lt;A href="https://www.jmp.com/support/help/en/19.1/#page/jmp/likelihood-aicc-and-bic.shtml#ww293087" target="_blank"&gt;Likelihood, AICc, and BIC)&lt;/A&gt;, agreement between models for terms inclusions, etc...&lt;/P&gt;
&lt;P&gt;Concerning your question about how to add/remove terms to build your models, the methods mentioned above can help. But more importantly, you need to define :&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Your objective&lt;/STRONG&gt;: Are you in a screening situation (where p-values may help to differentiate true signal from random variation or noise) ? Or in an optimization/prediction scenario (where accuracy of the model may matter more, so you may rely on RMSE and other predictive metrics)? Or you are in the middle, and not sure (Information criterion may help you compare models and determine the right balance for the terms and number of terms included by balancing accuracy with complexity).&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Your domain knowledge/expertise&lt;/STRONG&gt;: Are you able to remove/add some terms based on historical data or knowledge about the system you're studying ? Are there already some interactions or main effects known to be active and important ? Can you check/assess a model with domain experts to see if the model makes some sense ?&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Your validation strategy&lt;/STRONG&gt;: Once you have selected models that seem to be adequate for your responses, you need to validate these models. Based on your objective(s) and experimental budget, you may have different scenarii to make validation runs. Are you only interested in validating the optimum points ? Or do you want to validate the model on the factors ranges studied (and to do so, run combination of factors levels not in the design but respecting the factors ranges) ? There are no strict rules to validate your findings, but taking extra care (and time !) to run these validation experiments can greatly strenghten your confidence in the models (and the confidence from your reviewers !).&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Hope this answer will help you,&lt;/P&gt;</description>
      <pubDate>Mon, 21 Sep 2026 08:18:40 GMT</pubDate>
      <guid>https://community.jmp.com/t5/Discussions/How-to-select-model-terms-for-small-sample-size-SHould-I-use/m-p/972590#M110572</guid>
      <dc:creator>Victor_G</dc:creator>
      <dc:date>2026-09-21T08:18:40Z</dc:date>
    </item>
  </channel>
</rss>

