Here are my thoughts. Don't get lost in the specific tool per se. First gage accuracy can only be assessed through measurement of reference standards (e.g., NIST, CSA, TUV, BSI). I think your question is more likely about the entire measurement process and perhaps precision of the instrument. The hypothesis you state is the sample stack creation was contributing to variation (although you don't explain why). This is not the instrument, but the measurement process. What is that the technicians are doing differently? How did you arrive at the number 10 for your study? How were those 10 selected? What sources of sample (stack) variation will be represented in those 10 samples? Same questions for the 7 measures within stack. Why 7? Are these going to measured systematically or randomly? Same questions for technician. Why 3?
My recommendation is to treat this as a component of variation study. The components you want included in your study are technician, stack-to-stack, within stack. Your study does not appear to have the same technician measuring the same location on the same stack (measurement repeatability), but you could easily add that component. The study could be nested, systematic or crossed.
"All models are wrong, some are useful" G.E.P. Box